Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer does

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score)
  • Assigns actions to non-attendees
Case

v1.8 · synthetic · AI product, productivity

Usable with a quick edit9 of 9 checks passed

All constraints are met, the reasoning is sound, and the PRD is detailed enough for engineering and design leads to start work with only minor editorial polish.

Output

Huddle AI meeting summaries

Product requirements for engineering and design leads

Purpose and product decision

Huddle will generate a reviewable record of meeting decisions and actions from its transcripts. The host checks the draft, resolves ownership and publishes it. Approved actions can then be emailed to their owners.

Huddle serves 3,000 companies whose users currently write notes themselves. Speaker labels are wrong about 8% of the time, and roughly 15% of meetings include an uninvited participant. Calendar membership and transcript attribution therefore cannot establish attendance or ownership.

The MVP will not fulfil Sales’ request for immediate, unreviewed owner emails. It will generate drafts automatically after transcription and send emails automatically after explicit publication. Faster delivery does not justify assigning work to absent people. All requirements and numerical targets below are proposed product decisions.

Scope and intended outcome

Primary users are meeting hosts reviewing records and attendees receiving actions. Success means less time producing dependable notes, with no assignments to non-attendees.

MVP includes English transcripts, decisions, action extraction, attendance verification, host review, publication and owner emails. Exclude live summaries, recurring reminders, task-system integrations and automatic reassignment. Use existing meeting access and retention controls.

An action has one verified owner or remains unassigned. Split genuinely separate responsibilities into separate actions. Do not infer deadlines, commitments or decisions from tentative discussion.

Attendance is a release dependency

Engineering must first confirm whether Huddle retains authenticated participant identities and join/leave events. The supplied calendar list is insufficient. If these records are unavailable, build attendance instrumentation before enabling assignments; summaries may still launch with unassigned actions.

An eligible owner must have a stable person identifier linked to an authenticated, human participant session with a recorded join event for this meeting occurrence. Calendar invitees who never joined are ineligible. Bots and room devices are ineligible. Deduplicate reconnects by identity. Someone who left early remains eligible, although a later discussion does not prove they accepted work.

Uninvited users are eligible when their authenticated join is recorded. Guest display names, shared-room speaker labels and email addresses typed by a reviewer are not identity evidence. Unverified guests and people represented only by a room device remain ineligible in MVP; explain how they can join individually with verified identity for future meetings.

This deliberately sacrifices assignment coverage. The invariant is enforceable against verified participation records, not proof of physical human presence behind an account. If literal physical presence is required, that remains an unresolved product constraint and assignments must not launch under a stronger guarantee.

Ownership and extraction rules

The model extracts structured candidates with supporting transcript spans. Each decision contains its outcome and evidence timestamps. Each action contains a task, supporting spans, optional explicit due date and a nullable candidate participant identifier. Preserve the original date wording alongside any normalised date, using the meeting timezone; ambiguous dates stay unset.

Only explicit commitments or clearly agreed assignments qualify as actions. Suggestions, questions and rejected proposals do not. Where later discussion reverses a decision, present the final supported outcome. If the resolution is unclear, flag it for review instead of declaring a decision.

A model candidate is internal metadata, never an assigned owner. Server validation rejects identifiers outside the eligible attendance set before anything is displayed. The review UI starts every owner field as “Unassigned”; reviewers choose an eligible attendee after checking evidence. Do not preselect names from speaker labels, including for “I’ll do it”.

Render ownership only from validated structured fields. Generated task and decision wording must not embed owner claims that bypass these controls. Unsupported person-specific claims are withheld for review. A request that absent Alex should do something can become an unassigned “Confirm responsibility for…” follow-up only when that follow-up was actually agreed; otherwise omit it as an action.

Review and publication experience

The meeting page shows “Preparing summary” after the call, then “Draft needs review”. The host is the default reviewer; existing authorised meeting editors may also review. If no reviewer is available, retain the draft without sending owner emails.

Display separate Decisions and Actions sections. Each item links to its transcript passage and recording timestamp. Offer edit, delete and add controls. Actions show task, owner and due date. An attendance-only owner picker explains why invitees or unresolved guests cannot be selected. Provide a visible “Needs owner” state rather than silently dropping useful work.

Reviewers must confirm each chosen owner and can publish with unresolved actions. The publish dialogue shows recipients and unresolved-action count, with an explicit “Publish and email owners” control. Actions without owners remain visible but produce no email. Keyboard navigation, labelled controls and screen-reader status announcements are required.

Published records show their revision and approval time. Subsequent edits create a new revision. Ownership changes require the same validation and explicit publication, then notify affected owners with a correction. Never silently replace a published record with regenerated content.

Engineering contract and failure behaviour

Use meeting occurrence IDs, not recurring calendar series IDs. Store transcript version, attendance evidence version, model/prompt version, summary revision, item IDs, evidence spans, owner ID, approval actor/time and delivery status. Keep internal candidates separate from approved owner fields.

Process states are waiting for transcript, generating, needs review, published and failed. Trigger generation only after transcript finalisation. Transcript changes invalidate unpublished approvals and create a fresh draft; published revisions remain intact until a replacement is approved.

Enforce owner eligibility server-side on draft writes, publication, exports and notification dispatch. Manual edits and direct API calls cannot bypass it. If attendance evidence is missing, contradictory or unavailable, block ownership and sending rather than falling back to calendar data. Record rejection reasons for support without exposing transcript content in routine logs.

Use transactional publication and an outbox. Delivery keys combine meeting, published revision and recipient so retries do not duplicate emails. Email only approved actions owned by that recipient, with a permission-checked meeting link. Resolve addresses from verified identities, never model output. Block delivery if permissions or identity validation fail; show the reviewer the failure. Retry transient failures with backoff for up to 24 hours, then offer manual retry.

Treat transcripts as untrusted data: spoken instructions cannot change permissions, choose recipients or trigger tools. Generation has no email capability. Preserve tenant isolation and apply meeting deletion and retention policies to summaries and queued notifications.

Acceptance scenarios

  1. Invited Pat never joins but is named in an action: Pat is absent from the picker; API assignment is rejected; no owner email is sent.
  1. Uninvited authenticated Lee joins: Lee is eligible and can receive an action after reviewer confirmation and publication.
  1. Several people join through one room device: no individual becomes eligible from its label; actions remain unassigned unless independent authenticated participation exists.
  1. Overlapping speech mislabels a commitment: no owner is preselected; the reviewer can inspect evidence and choose a verified attendee or leave it unassigned.
  1. No decisions or actions are supported: show explicit empty states rather than inventing content. Failed or incomplete transcription shows a retry state and sends nothing.
  1. A forged owner ID, missing attendance record or duplicate publish event cannot create an invalid assignment or duplicate email. Revoked access blocks queued delivery.

Evaluation and release

Build a consented, de-identified evaluation set of at least 200 meetings, deliberately including overlap, shared rooms, absent invitees, uninvited guests, duplicate names, reversals and ambiguous commitments. Human annotators label attendance, final decisions, actions, ownership and evidence; adjudicate disagreements independently. Report each difficult slice separately.

Release gates: zero non-attendee assignments or deliveries across end-to-end and adversarial tests; at least 95% precision and 85% recall for supported decisions and actions, measured separately. Zero observed violations is necessary but insufficient: eligibility enforcement must also pass API and notification-path tests. Model confidence alone never permits assignment.

Pilot with ten opt-in companies before staged rollout. Target draft availability within two minutes of final transcript for 95% of meetings up to 60 minutes, and a 50% reduction in median note-preparation time against pilot baseline. Track owner corrections, unassigned-action rate, publication rate and delivery failures. Any non-attendee assignment disables assignments and sending pending investigation. Expand only after two weeks without a critical violation and acceptable review effort; do not reduce abstention merely to improve coverage.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll factual claims about the current situation come directly from the brief or context, with no invented numbers or false statements.
  • Addresses the actual decisionThe output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
  • Respects explicit constraintsIt respects the never‑assign‑to‑non‑attendees constraint through eligibility checks, host review, and server‑side enforcement; length, reader and form are all within the brief’s requirements.
  • Identifies material uncertaintyIt identifies the unknown availability of authenticated join/leave records, bounds the impact, and states that if missing, assignments must be blocked until attendance instrumentation is built.
  • Avoids unsupported claimsInterpretations are clearly labelled as product decisions or derived from the supplied evidence; no confident claim goes beyond what the evidence supports.
  • Produces the required deliverableThe output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
  • Specifies AI failure behaviourThe PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
  • Success settles trade-offsSuccess is defined with measurable targets (zero non‑attendee assignments, 95%/85% precision/recall, 50% reduction in note time) and an explicit trade‑off rule (sacrificing coverage to preserve correctness).
  • Proposes tests that could failEvery acceptance scenario, release gate, and pilot has numeric thresholds, a clear measurement window, and a triggered action (e.g., disable assignments if any non‑attendee assignment occurs).

Grades and run details

Decision model 94 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly11%
  • passAddresses the actual decision98%
  • passRespects explicit constraints67%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims22%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs71%
  • partialProposes tests that could fail68%
Run
Run
#1
Time to output
3.2 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD