Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer does

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score)
  • Assigns actions to non-attendees
Case

v1.8 · synthetic · AI product, productivity

Usable with a quick edit5 of 9 checks passed

Yes. A PM could quickly add measurement windows to the metrics and rephrase the unsupported problem-statement claims without redoing any structural or reasoning work.

Output

PRD: AI Meeting Summaries and Action Items

Product: Huddle Audience: Engineering and Design leads Status: Proposed MVP Launch decision: Ship reviewed summaries first; defer immediate, unreviewed owner emails.

1. Problem and opportunity

Huddle serves 3,000 companies. Meetings are recorded and transcribed with speaker labels, but attendees currently write their own notes afterward. This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

We will generate a concise meeting summary containing decisions and action items, with evidence linking each item to the recording.

Two known limitations shape the design:

  • Speaker labels are wrong approximately 8% of the time, especially during overlapping speech and shared-room participation.
  • Approximately 15% of meetings include someone absent from the calendar invite.

Therefore, neither speaker labels nor calendar invitations are sufficient evidence for assigning work.

Non-negotiable requirement: An action must never be assigned to someone who did not attend the meeting. When attendance or identity cannot be established, leave the action unassigned.

2. Goals and non-goals

Goals

  • Reduce time spent producing post-meeting notes.
  • Capture decisions accurately, distinguishing decisions from proposals.
  • Extract actionable commitments with owners and deadlines when supported.
  • Include verified, uninvited attendees without treating all invitees as attendees.
  • Make every decision and action inspectable and correctable.
  • Deliver owner notifications only after safe assignment and publication.

Non-goals for MVP

  • Live summaries or action assignment during meetings.
  • Assigning work to absent colleagues, even if requested during the meeting.
  • Inferring deadlines or assigning owners based on job title.
  • External task-system integrations.
  • Automatically publishing unreviewed summaries.
  • Resolving identity using voice biometrics.

3. Users and primary workflow

Reviewer: Meeting organizer, or an explicitly designated attendee. Reviews and publishes the summary.

Attendee: Reads the summary, follows evidence links, and reports corrections.

Action owner: Receives published actions assigned to them.

Workflow:

  1. Meeting ends; Huddle completes transcript processing.
  2. Huddle reconciles actual attendance and generates a private draft.
  3. Reviewer checks decisions, unresolved ownership, and attendance exceptions.
  4. Reviewer edits and publishes.
  5. Huddle makes the summary available to authorized viewers and, where enabled, emails owners their approved actions.

If no eligible reviewer is available, the draft remains unpublished. A non-attending organizer may designate an attending reviewer but cannot resolve attendance exceptions themselves.

4. Attendance and identity requirements

Engineering must establish an attendance ledger separate from the calendar invite.

Each record contains:

  • Meeting ID and stable person ID, where resolved.
  • Join/leave intervals and participant-session IDs.
  • Identity source: authenticated join, verified guest identity, or shared-room roster.
  • Verification state and evidence.
  • Reviewer confirmation details, where required.

Calendar invitees are suggestions for identity matching, not proof of attendance. A directory search result is also not attendance evidence.

Assignment eligibility

An owner must map to a unique person with verified attendance in that meeting:

  • Authenticated individual join: Eligible when attendance telemetry records the person’s session.
  • Guest join: Eligible only after identity is resolved and verified; a display name alone is insufficient.
  • Shared meeting room: The room endpoint proves room participation, not individual presence. A participating reviewer must explicitly confirm each person physically present before they become eligible.
  • Uninvited attendee: Eligible through the same verification rules; invitation status is irrelevant.
  • Invited but absent person: Ineligible.

If telemetry is unavailable, identity is ambiguous, or attendance is disputed, block assignment. Keep the action unassigned with an explanation.

The owner picker must show only eligible attendees. Every write path—including generation, manual edits, APIs, and publication—must enforce eligibility server-side. Model confidence must never override this rule.

These controls cannot independently observe who is physically in a room. If trustworthy roster confirmation is unavailable, shared-room individuals remain ineligible rather than weakening the constraint.

5. Summary and extraction requirements

The draft has four sections:

  1. Overview: Two to four sentences describing the meeting’s purpose and outcome.
  2. Decisions: Clear statements of agreed outcomes.
  3. Action items: Task, owner or “Unassigned,” deadline if explicit, and status.
  4. Needs review: Uncertain decisions, ownership, identity, or attendance.

Each decision and action must include transcript-span references and recording timestamps. Preserve the original supporting text even when the reviewer edits the final wording.

Decisions

Capture explicit agreement or a clearly accepted conclusion. Do not convert brainstorming, recommendations, or unresolved debate into decisions. If discussion reverses an earlier decision, represent the final outcome and retain supporting evidence.

Actions

Extract a concrete deliverable or next step. Record a deadline only when explicitly stated; normalize relative dates using the meeting’s date and timezone.

Examples:

  • “I’ll send the revised deck tomorrow.” → Task with a candidate owner, subject to attribution review.
  • “Someone should update the deck.” → Unassigned action.
  • “Ask Priya to update it,” when Priya was absent → Unassigned action; note that an absent person was mentioned.
  • “We could revisit this next quarter.” → Not an action without a clear commitment.

Speaker labels provide evidence, not authority. Overlap, shared-room speech, ambiguous pronouns, and conflicting commitments must trigger ownership review. Candidate owners are draft suggestions, not assignments, until approved.

6. Review experience

The summary page opens in Draft state with a clear “Review before sharing” banner.

Design must provide:

  • Editable decision and action cards.
  • Evidence links opening the relevant transcript and recording segment.
  • Visible distinction between verified attendance and uncertain attribution.
  • “Unassigned” as a first-class owner option.
  • A dedicated queue for identity and attendance exceptions.
  • A publish preview showing recipients and notification settings.

For each assigned action, the reviewer must approve the owner. Bulk approval is permitted for straightforward items, but flagged attribution requires item-level review.

Attendance verification and ownership approval are separate controls: confirming that someone attended does not prove they accepted a task.

Publication is blocked if any assigned owner is ineligible. Unassigned actions and unresolved decision candidates do not block publication; uncertain decisions must be omitted from the final Decisions section or explicitly labeled unresolved.

7. Email policy and Sales request

Sales requests automatic emails immediately after each meeting. MVP will not send unreviewed action assignments. Given the known speaker errors, immediate delivery would spread incorrect ownership before correction.

Instead:

  • Notify the reviewer when the draft is ready.
  • Default owner emails to off; enable through workspace settings.
  • When enabled, automatically email approved owners after publication.
  • Send each owner only their actions, meeting title, deadlines, and an authorized summary link.
  • Never email a calendar invitee solely because they were invited.
  • Do not send to guests without a verified address and permitted access.

Email dispatch must recheck eligibility and publication version. Use idempotency keys to prevent duplicate sends. Published ownership changes generate a correction for the previous recipient and an updated notification for the new eligible owner.

Future immediate automation requires a separate launch decision and evidence that it preserves the attendance constraint.

8. Technical implementation

Use an asynchronous pipeline:

`Meeting ended → Transcript ready → Attendance reconciled → Draft generated → Reviewed → Published → Notifications`

Proposed records:

  • `Summary`: meeting ID, transcript version, state, generated content, revision.
  • `Decision`: text, evidence spans, review state.
  • `Action`: task, proposed owner, approved owner, deadline, evidence, review state.
  • `AttendanceRecord`: identity, evidence, verification state.
  • `AuditEvent`: actor, change, timestamp, revision.

Generation returns schema-validated structured data. Treat transcript content as untrusted input: instructions spoken during a meeting must not alter system rules, access controls, or recipient selection.

Owner writes and publication run through a centralized eligibility validator. Notification jobs check the current revision before sending. Attendance corrections invalidate affected assignments and cancel queued notifications. Maintain an audit trail rather than silently rewriting history.

9. Reliability, privacy, and access

Target draft availability: within five minutes of transcript completion for 95% of supported meetings.

Generation failures must leave recordings and transcripts usable, show retry status, and send no owner emails.

Drafts are visible only to reviewers. Published summaries inherit meeting-content access rules; attendance alone does not grant access. Recording evidence links must enforce existing permissions. Apply workspace retention and deletion policies to summaries, attendance evidence, and generated artifacts.

10. Success metrics and acceptance

Proposed launch gates

  • Zero assignments to ineligible attendees across automated invariant tests.
  • Zero emails to ineligible owners in the pilot; any confirmed incident pauses rollout.
  • At least 90% of published decisions require no substantive reviewer correction.
  • Median review time below three minutes for meetings under one hour.
  • At least 30% less self-reported note-taking time in the pilot.

Measure owner-correction rate separately from attendance violations: a wrong attending owner is still a quality failure.

Required acceptance cases

Test absent invitees, verified uninvited guests, shared rooms, overlapping speakers, duplicate names, failed identity resolution, missing telemetry, absent people named in requests, transcript reprocessing, attendance revoked before dispatch, and repeated notification retries.

In every uncertain-attendance case, expected behavior is an unassigned action and no owner email.

11. Rollout and dependencies

Start with internal meetings, then an opt-in pilot across varied company sizes and shared-room usage. Expand only after launch gates pass.

Before implementation, confirm availability of participant-session telemetry, guest identity verification, room-roster confirmation, reviewer permissions, and verified email addresses. If these dependencies are missing, ship summaries with unassigned actions for affected participants rather than substituting calendar membership as attendance proof.

Check by check

Got wrong · 4
  • Uses the supplied evidence correctlyThe output presents 'This creates duplicated effort, inconsistent records of decisions, and missed follow-through' as a current-situation fact, but it is not in the supplied context and does not follow by arithmetic.
  • Avoids unsupported claimsThe output states 'This creates duplicated effort...' and 'immediate delivery would spread incorrect ownership' as established facts without labelling them as inferences or forecasts, although the supplied context only suggests these outcomes.
  • Success settles trade-offsThe success metrics have numeric targets but no explicit time frame (e.g., pilot duration), and the output lacks a single clear trade-off rule stating which goal wins when precision and coverage conflict, though unassigned-over-misassignment is implied.
  • Proposes tests that could failThe launch gates and acceptance cases have numeric thresholds and some actions, but they lack measurement windows long enough to read out the effects, and not every gate specifies the action if the threshold is not met.
Got right · 5
  • Addresses the actual decisionThe output commits early to 'Ship reviewed summaries first; defer immediate, unreviewed owner emails' and states that future immediate automation requires evidence that it preserves the attendance constraint.
  • Respects explicit constraintsThe output enforces the non-assignment constraint through an attendance ledger, eligibility checks, server-side validation, and reviewer approval; it also addresses the Sales request by deferring auto-email until after review.
  • Identifies material uncertaintyThe output names specific unknowns such as shared-room physical attendance, missing telemetry, and ambiguous identity, and says how they would be resolved (reviewer confirmation, unassigned actions) and what would change the call (evidence for future automation).
  • Produces the required deliverableThe output is a complete PRD aimed at engineering and design leads, covering decisions, actions, attribution, editing, and failure behaviour, and it appears to be within the requested 1,000-1,500 word range.
  • Specifies AI failure behaviourThe PRD explicitly specifies what happens when identity or attendance is uncertain (unassigned actions, no emails), when reviewers are unavailable (draft remains unpublished), and when generation fails (no owner emails, recordings usable).

Claims the judge couldn’t find in the brief

  • This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly5%
  • passAddresses the actual decision99%
  • passRespects explicit constraints70%
  • passIdentifies material uncertainty87%
  • failAvoids unsupported claims20%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • partialSuccess settles trade-offs54%
  • partialProposes tests that could fail76%
Run
Run
#1
API response time
47 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD