Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer does

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score)
  • Assigns actions to non-attendees
Case

v1.8 · synthetic · AI product, productivity

Needs more than a quick edit5 of 9 checks passed

It includes an unsupported claim about note quality that a product manager would need to verify, and it omits a time frame for its success metrics, requiring more than a quick edit.

Output

PRD: AI Meeting Summaries — Decisions & Action Items

Sep 24, 2026 · @Martin Slaney

Summary

We will generate a summary after every recorded Huddle meeting that lists the decisions made and the action items, each with an owner who was verifiably in the meeting. The organizer reviews the summary and sends it in one click; nothing is emailed without that step in v1.

Problem. Today attendees write their own notes after the call. Notes are inconsistent, often missing, and action items lose their owners within a day. We already have what we need to fix this: a recording, a transcript with speaker labels, and the calendar attendee list.

Why this is not trivial. Two of those inputs are unreliable. Speaker labels are wrong about 8% of the time, mostly during crosstalk and when several people share a meeting-room device. About 15% of meetings include someone who joined without a calendar invite. A naive build would misassign owners often enough to lose trust in the first week.

Goals and metrics

The goal is summaries people send without rewriting, with zero actions assigned to non-attendees.

Goals

  • Replace manual post-meeting notes for recorded meetings.
  • Every action item has a correct owner or is explicitly marked unassigned.
  • Get action items to owners within minutes of the meeting ending.

Non-goals (v1)

  • Pushing tasks into Jira, Asana or Linear.
  • Summaries for unrecorded meetings or live, in-call summaries.
  • Improving diarization itself. We design around the 8% error rate, not fix it.

Success metrics

MetricTargetType
Actions assigned to someone not in the meeting0Hard constraint, blocks release
Owner precision on items sent with an owner≥ 97%Quality
Summaries sent within 30 min of meeting end≥ 60% of recorded meetingsAdoption
Summaries sent with no edits≥ 50%Quality
Owner reassignments per sent summary≤ 0.3Quality

Owner precision is measured from organizer edits plus a weekly labelled sample of 200 meetings.

Users and core flow

The organizer owns the summary; attendees receive it.

  • Organizer: reviews, fixes owners, sends. Usually the calendar invite's creator; can hand off to any attendee.
  • Action owner: receives their items by email and in the app.
  • Other attendees: receive the full summary once sent.

```mermaid flowchart LR A[Meeting ends] --> B[Transcript +<br/>join log ready] B --> C[Build verified<br/>roster] C --> D[Extract decisions<br/>and actions] D --> E[Roster filter +<br/>confidence check] E --> F[Organizer review] F --> G[Send to owners<br/>and attendees] ```

The roster filter runs as deterministic code after the model, never inside the prompt alone.2 Target: summary ready for review within 5 minutes of meeting end for meetings up to 90 minutes.

Requirements

An owner can only be chosen from the verified roster: people who actually joined, not people who were invited.

Verified roster

1. R1. Build the roster from the join log, not the calendar invite.1 Invited people who never joined are excluded. 2. R2. Each roster entry has an identity tier: - Verified: joined signed in to a Huddle account, or by a dial-in matched to a known number. - Guest: joined with a display name only (most of the 15% uninvited joiners). - Room: a meeting-room device; people behind it are unknown individuals. 3. R3. Match uninvited joiners to workspace accounts by signed-in account only. Never match on display name alone.

Summary content

  1. R4. Decisions: a short statement of each decision, with a transcript timestamp link. No owner field.
  2. R5. Action items: task, owner (or "Unassigned"), due date if said aloud, timestamp link, and confidence.
  3. R6. Every item links to its evidence in the transcript so reviewers can check it in one click.

Assignment rules

  1. R7. The model returns a proposed owner per item. Code then checks it against the roster. Any owner not on the roster is replaced with "Unassigned", and the named person is shown as a mention ("mentions: Dave, not in meeting").
  2. R8. Owner comes from content first: a named request ("Priya, can you…") beats the speaker label of "I'll do it".
  3. R9. Mark an owner low confidence when the evidence depends on a speaker label in an overlapping segment, a Room entry, or a Guest entry.
  4. R10. Low-confidence items show as "Suggested: Priya" and need organizer confirmation before any email goes out.
  5. R11. Room and Guest entries cannot receive email. The organizer can reassign to a verified person or leave it unassigned.
  6. R12. Log every model proposal, filter result and organizer edit. This is our precision measurement.

Handling the known limits

Both limits hit ownership, not decisions, so we contain them at the owner step.

LimitWhat goes wrongHow we handle it
Speaker labels wrong \~8% (crosstalk)"I'll take that" credited to the wrong personContent-first owner rule (R8); overlap segments flagged low confidence (R9)
Speaker labels wrong \~8% (shared room)One label covers several peopleRoom tier; items resolved to a room need a named person or stay unassigned (R11)
\~15% of meetings have uninvited joinersReal participant missing from roster, or matched to the wrong accountRoster from join log (R1); account match only (R3); Guests can own items but get no email
Invited but absentModel assigns to a name from the inviteExcluded from roster; hard filter (R7)

The error compounds per meeting. If each item's owner is wrong 8% of the time, a meeting with 5 action items has a 34% chance of at least one wrong owner (1 − 0.92⁵). That is why confirmation sits before email, not after.

Sales request: auto-email on meeting end

Decision: not in v1 as requested. We ship one-click organizer send now, and gated auto-send in phase 2 once measured precision earns it.

Why not now. An email to the wrong owner is the most visible failure this feature can have. With the error rate above, roughly a third of 5-item meetings would email someone a task that isn't theirs. With external guests, it may land outside the customer's company. Early trust is hard to win back.

What Sales gets instead.

  • v1: the organizer gets a "Summary ready" notification within 5 minutes and sends with one click. Goal: sent within 30 minutes (see metrics).
  • Phase 2: a workspace setting to auto-send at meeting end, but only items that are high confidence and owned by a Verified attendee. Everything else still waits for the organizer. It ships only when owner precision on high-confidence items is ≥ 97% for 4 consecutive weeks.
  • Auto-send never emails Guest or Room entries, or anyone outside the customer's domain, unless an admin enables it.

Design notes

The review screen must make checking owners faster than writing notes.

  • Decisions first, then action items grouped by owner, then Unassigned.
  • Low-confidence owners look visibly different ("Suggested") and are editable inline from a roster picker, never free text.
  • Each item plays its transcript clip on click.
  • The Send button states who gets email: "Send to 4 owners and 7 attendees".
  • Empty state: if no decisions or actions are found, say so plainly rather than inventing any.

Rollout

PhaseScopeExit criteria
InternalHuddle's own meetingsZero non-attendee assignments; precision logged
Beta50 opted-in companies≥ 97% owner precision on sent items; ≥ 40% sent within 30 min
GAAll 3,000 companies, organizer sendBeta targets held for 4 weeks
Phase 2Gated auto-send settingSee Sales section

Open questions

  • [ ] Is the join log's signed-in identity reliable enough to call "Verified"? Engineering to confirm.
  • [ ] Who organizes when the invite creator didn't attend? Proposed: first Verified attendee from the invite.
  • [ ] Do customers need to disable summaries per meeting for sensitive calls? Legal and Sales to confirm.
  • [ ] Retention: how long do we keep drafts and the evidence log?

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    Source: Scenariothe attendee list comes from the calendar invite

    What we’d changeTreat the join log as a dependency to confirm, not an input that exists today. Say what proves attendance, and don't promise zero non-attendee assignments until that's settled.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: ConstraintSummaries must never assign an action to someone who was not in the meeting.

    What we’d changeApply the roster check to organiser edits and the final send as well, not only to the model's output.

Check by check

Got wrong · 3
  • Uses the supplied evidence correctlyThe output claims that today's notes are inconsistent, often missing, and action items lose owners within a day; this is not in the supplied context and invents current performance.
  • Avoids unsupported claimsThe statement about notes being inconsistent and action items losing owners is presented as fact without the supplied evidence supporting it.
  • Proposes tests that could failInternal and beta exit criteria lack a measurement window (beta has no duration, internal only 'zero non‑attendee assignments' without a window), and not all phases state the action each result triggers.
Mixed · 1
  • Success settles trade-offsSuccess metrics lack a time frame, and no explicit trade-off rule is stated beyond the hard constraint, so the PRD does not define which goal wins when two conflict.The two graders disagreed on this one.
Got right · 5
  • Addresses the actual decisionIt commits early to organizer-reviewed summaries with no auto-email in v1, says what would change that (precision data), and frames the choice for engineering and design leads.
  • Respects explicit constraintsThe PRD enforces the zero-non-attendee constraint with a verified roster and hard filter, and the output is within the 1,000–1,500 word aim.
  • Identifies material uncertaintyIt names specific unknowns (join log reliability, organizer when creator absent, per-meeting disable, retention) and how they would be resolved.
  • Produces the required deliverableThe output is a PRD that gives requirements, flow, design notes, and rollout; engineering and design leads could start work from it with light edits.
  • Specifies AI failure behaviourIt specifies low‑confidence flags, unassigned fallback, and organizer confirmation before email, covering AI uncertainty and errors.

Claims the judge couldn’t find in the brief

  • Notes are inconsistent, often missing, and action items lose their owners within a day.

Grades and run details

Decision model 72 · LLM judge 5 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly47%
  • passAddresses the actual decision99%
  • passRespects explicit constraints67%
  • passIdentifies material uncertainty86%
  • failAvoids unsupported claims36%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs83%
  • partialProposes tests that could fail66%
Artefacts
Run
Run
#1
Time to output
80 s
Submitted
24 Sept 2026

Needs more than a quick edit3 of 9 checks passed

The PRD lacks essential trade-off rules and rollout gates a PM would need. It would require rework to define what success rejection looks like and how to make go/no-go decisions.

Output

Product Requirements Document (PRD): AI-Generated Meeting Summaries & Action Items

1. Document Overview & Objective

1.1 Product Name

Huddle Intelligent Summaries (Internal Codename: Project Cue)

1.2 Target Audience

Engineering Leads, Frontend/Backend Engineers, and Product Designers.

1.3 Objective

Build an automated, post-meeting intelligence feature for Huddle that processes meeting transcripts to extract structured decisions, generate comprehensive summaries, and assign action items to attendees. The feature must account for real-world messy transcript data (speaker errors, uninvited participants) and adhere to strict safety boundaries regarding action item assignments and automated distribution.

1.4 Background & Scenario

Huddle is a video-meeting application currently utilized by 3,000 enterprise and mid-market companies. * Existing Infrastructure: Meetings are automatically recorded and transcribed with basic speaker labeling. Attendee lists are reliably synchronized from the originating calendar invite. * Current State & Pain Point: Users currently rely on manual post-meeting note-taking, leading to dropped tasks, misaligned follow-ups, and fragmented alignment. * Value Proposition: Automating meeting wrap-ups directly inside Huddle saves hours of manual administrative overhead per week4, driving higher platform stickiness and daily active use.

---

2. Core User Stories & Personas

2.1 Personas

  • The Engineering Manager (User): Attends 5–6 back-to-back technical syncs daily. Needs clear, unequivocal records of architectural decisions and precise engineering follow-ups without digging through raw, messy transcripts.
  • The Account Executive / Sales Rep (User): Operates on rapid customer cadences. Requires immediate, zero-friction distribution of post-meeting action items to stakeholders without manual typing.
  • The Engineering & Design Leads (Builders): Require explicit edge-case handling, system constraints, deterministic UI layouts, and scalable pipeline specs to build and ship the feature efficiently.

2.2 Key User Stories

  1. As a meeting participant, I want Huddle to automatically generate a structured summary and action item list the moment a meeting concludes, so I don't have to spend time writing manual recap notes.
  2. As an attendee with assigned tasks, I want clear ownership labels on my action items so there is never any ambiguity about who is driving a deliverable.
  3. As a sales representative, I want my action items emailed to me automatically as soon as the meeting ends, so I can immediately execute client follow-ups.
  4. As a meeting host, I want to review, edit, or override incorrect speaker attributions or bad action assignments before summaries are locked or shared externally.

---

3. Detailed Functional Requirements

3.1 Data Inputs & Pre-Processing Pipeline

The AI summarization engine triggers asynchronously via webhook immediately upon meeting termination (`recording.ended`).

  • Input A: Transcript with Speaker Labels: Full text payload containing timestamped utterances, text strings, and speaker IDs. Known limitation: Speaker labels suffer from an 8% error rate, primarily caused by cross-talk, overlapping dialogue, or multi-person huddle rooms.
  • Input B: Calendar Invite Attendee List: Ground-truth list of officially invited participants (Email, Name, Role).1 Known limitation: 15% of meetings include an uninvited participant (e.g., ad-hoc joiners via shared link).
  • Input C: Meeting Metadata: Duration, title, organization ID, and recording URL.

3.2 Summarization & Extraction Logic (LLM Pipeline)

The backend processing pipeline must execute three logical passes: 1. Meeting Overview Generation: Condense the transcript into a 3–5 sentence executive summary alongside 3 key thematic highlights. 2. Decision Extraction: Isolate explicit agreements, architectural choices, or policy resolutions made during the call. Format as an bulleted list of declarative statements (e.g., "Decided to migrate the primary auth service to OAuth2 by Q3"). 3. Action Item Extraction & Assignment Engine: * Extract explicit tasks, owners, and deadlines. * Crucial Constraint (Safety Guardrail): The engine must never assign an action item to someone who was not present in the meeting. * Resolution Strategy for Uninvited Attendees: If an uninvited participant (who joined without a calendar invite) is identified as the owner of a task, the system must cross-reference their presence against the audio/transcript participant logs. If verified present, they are added to a dynamic "Extended Participants" list for assignment validation. If their identity cannot be verified or they were absent, the task defaults to unassigned or falls back to the meeting host with a UI warning flag2.

3.3 Sales Request: Automated Instant Emailing

  • Requirement: Sales teams have requested that action items be emailed automatically to their owners the second a meeting ends.
  • Product Specification & Guardrail: Due to the 8% speaker label error rate and potential hallucination risks, fully autonomous external emailing on meeting close creates high business risk.
  • Implementation Design:
  • Implement an Opt-In Workspace Policy configured by organization admins.
  • For workspaces with "Auto-Send Sales Recaps" enabled, summaries and assigned action items are dispatched via email to internal organization members immediately upon processing completion.3
  • For external participants or non-opt-in workspaces, a 10-minute "Grace Period / Review Window" notification banner appears in the Huddle desktop client, allowing hosts to review and edit assignments before dispatch.

---

4. User Experience & Interface (UX/UI) Specifications

4.1 Post-Meeting View (Huddle Web & Desktop App)

Upon meeting conclusion, a new tab labeled "AI Summary" appears alongside the existing "Transcript" and "Recording" tabs within the meeting details drawer.

``` +-------------------------------------------------------------------+ | Huddle Meeting: Q3 Architecture Sync [Share] [Export]| +-------------------------------------------------------------------+ | Tabs: [Overview] [Transcript] [Recording] | |-------------------------------------------------------------------| | EXECUTIVE SUMMARY | | The team reviewed the Q3 scaling bottlenecks. Consensus was | | reached on moving auth infrastructure to OAuth2. Database latency | | issues will be mitigated via Redis caching layers. | | | | KEY DECISIONS | | • Migrate core auth service to OAuth2 by August 15. [Edit] | | • Adopt Redis cluster for session state caching. [Edit] | | | | ACTION ITEMS | | [x] Investigate OAuth2 token expiry edge cases | | Owner: Sarah Jenkins (Calendar Invite) | Due: Next Tuesday | | [ ] Set up staging Redis cluster | | Owner: Unassigned ⚠️ (Verify Attendee) | Due: Friday | | | | [Send Sales Recap Email Now] [Approve & Lock] | +-------------------------------------------------------------------+ ```

4.2 Interactive States & Edge-Case UI Handling

  • Unassigned / Unverified Owner State: If the model detects an action item owner who was an uninvited participant or has an ambiguous speaker label, render a yellow warning badge: `⚠️ Unverified Owner — Click to Reassign`.
  • Inline Editing: Every decision and action item card features an inline pencil icon allowing hosts or assignees to edit text, change owners via a searchable dropdown of confirmed attendees, or delete hallucinated items.

---

5. Technical Architecture & Engineering Guidelines

5.1 Pipeline Flow & Data Schema

  1. Webhook Listener: Receives `meeting.ended` event containing metadata and storage URIs.
  2. Context Aggregator: Merges calendar invite list (`invitedusers[]`) with actual audio track speaker diarization outputs (`detectedspeakers[]`).
  3. LLM Structured Extraction: Prompt engineering optimized for JSON output schema enforcement (using strict JSON mode / function calling):

```json { "meetingid": "string", "summary": "string", "decisions": ["string"], "actionitems": [ { "task": "string", "assigneeemail": "string | null", "assigneename": "string", "isverifiedattendee": boolean, "due_date": "ISO-8601 | null" } ] } ```

5.2 Error Handling & Fallbacks

  • Speaker Label Inaccuracy (8% Error Rate): Mitigation via UI feedback loops. If an assignee name has low confidence score (< 0.85) from the diarization model, flag the action item with a review indicator in the UI.
  • Uninvited Participant Handling (15% Occurrences): The validation layer must query the roster of active audio streams. If an email is missing from the calendar invite but present in audio metadata, tag them as `guest_participant` and permit assignment only if their voiceprint or explicit self-identification ("Hi, this is Dave from Finance joining late") is logged in the transcript text. Otherwise, fail-safe to unassigned.

---

6. Non-Functional Requirements & Constraints

  • Latency: Summary generation and database write must complete within < 120 seconds of meeting termination for a standard 45-minute meeting.
  • Security & Privacy: Meeting transcripts and summaries must be encrypted at rest (AES-256) and in transit (TLS 1.3). Customer data must never be used to train foundational models across tenant boundaries.
  • Reliability: 99.9% pipeline uptime. In the event of an LLM timeout or parsing failure, the system must gracefully fall back to displaying raw transcripts with an error banner stating "AI summary generation delayed. Click to retry."

---

7. Metrics & Success Criteria (KPIs)

  1. Adoption Rate: % of completed meetings where users view, edit, or export the AI summary (Target: > 40% within 30 days post-launch).
  2. Action Item Completion Rate: Increase in marked-complete action items compared to legacy manual workflows.
  3. Accuracy & Correction Rate: Frequency of manual edits made to action item owners and speaker assignments (Target: < 5% manual override rate, validating low false-positive rates for uninvited attendee checks).
  4. Sales Workflow Efficiency: Reduction in post-call administrative turnaround time reported by enterprise sales teams using the automated email dispatch feature.

What a PM had to fix

  1. 1Constraint missedRestore the constraintStart again

    Source: ConstraintSummaries must never assign an action to someone who was not in the meeting.

    What we’d changeBuild the list of possible owners from who actually joined, not who was invited, and enforce it in code on generated owners, manual edits and sending. A model-set is_verified_attendee flag can't guarantee it.

  2. 2Constraint missedRestore the constraintSubstantial rework

    What we’d changeLeave items with an unverifiable owner unassigned. Handing them to the host invents responsibility nobody agreed to.

  3. 3Constraint missedRestore the constraintSubstantial rework

    Source: Known limitsSpeaker labels are wrong about 8% of the time

    What we’d changeDon't auto-send. With speaker labels wrong 8% of the time, action items would reach the wrong owners before anyone checks them. Keep organiser review before any email.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent the time saved as a goal to measure, not a known benefit.

Check by check

Got wrong · 5
  • Uses the supplied evidence correctlyThe output claims Huddle is used by 'enterprise and mid-market companies' and has a 'Huddle Web & Desktop App', neither of which is in the supplied context.
  • Identifies material uncertaintyIt mentions error rates but does not specify what conditions would change the design decision (e.g., an override-rate threshold that would trigger disabling auto-email).
  • Avoids unsupported claimsIt presents 'enterprise and mid-market' and the existence of a desktop client as established facts without backing from the supplied evidence.
  • Success settles trade-offsNo explicit trade-off rule is given (e.g., accepting lower coverage to keep assignment precision above a stated level), and not all success metrics have targets and time frames.
  • Proposes tests that could failNo tests, gates, or kill criteria with numeric thresholds, measurement windows, and consequents are proposed.
Mixed · 1
  • Respects explicit constraintsThe constraint that summaries never assign actions to non-attendees is enforced via verification and fallback to unassigned/host; the PRD also fits the requested form and length.The two graders disagreed on this one.
Got right · 3
  • Addresses the actual decisionThe PRD commits to a clear design, addressing the Sales request with an opt-in auto-email and a review window, and is unambiguous for engineering and design leads.
  • Produces the required deliverableThe output is a structured PRD with functional specs, UI mockups, and technical architecture, within 1,000–1,500 words, and is usable by engineering and design leads.
  • Specifies AI failure behaviourIt defines UI warning states for unverified owners, low confidence, and fallback to raw transcript on LLM timeout.

Claims the judge couldn’t find in the brief

  • Huddle is used by 3,000 enterprise and mid-market companies.
  • Huddle has a Huddle Web & Desktop App.

Grades and run details

Decision model 50 · LLM judge 4 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly56%
  • passAddresses the actual decision42%
  • failRespects explicit constraints40%
  • partialIdentifies material uncertainty27%
  • failAvoids unsupported claims44%
  • passProduces the required deliverable59%
  • passSpecifies AI failure behaviour95%
  • partialSuccess settles trade-offs54%
  • partialProposes tests that could fail54%
Artefacts
Run
Run
#1
Time to output
15 s
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD