Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given4 items: Scenario, Known limits, Constraint, Request from Sales
ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer doesThe answer key the graders mark against

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score at 40)
  • Assigns actions to non-attendees
Case

Mid-level · v1.10 · synthetic · AI product, productivity

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit5 of 9 checks passed

Yes. A PM could quickly add measurement windows to the metrics and rephrase the unsupported problem-statement claims without redoing any structural or reasoning work.

Output

PRD: AI Meeting Summaries and Action Items

Product: Huddle Audience: Engineering and Design leads Status: Proposed MVP Launch decision: Ship reviewed summaries first; defer immediate, unreviewed owner emails.

1. Problem and opportunity

Huddle serves 3,000 companies. Meetings are recorded and transcribed with speaker labels, but attendees currently write their own notes afterward. This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

We will generate a concise meeting summary containing decisions and action items, with evidence linking each item to the recording.

Two known limitations shape the design:

  • Speaker labels are wrong approximately 8% of the time, especially during overlapping speech and shared-room participation.
  • Approximately 15% of meetings include someone absent from the calendar invite.

Therefore, neither speaker labels nor calendar invitations are sufficient evidence for assigning work.

Non-negotiable requirement: An action must never be assigned to someone who did not attend the meeting. When attendance or identity cannot be established, leave the action unassigned.

2. Goals and non-goals

Goals

  • Reduce time spent producing post-meeting notes.
  • Capture decisions accurately, distinguishing decisions from proposals.
  • Extract actionable commitments with owners and deadlines when supported.
  • Include verified, uninvited attendees without treating all invitees as attendees.
  • Make every decision and action inspectable and correctable.
  • Deliver owner notifications only after safe assignment and publication.

Non-goals for MVP

  • Live summaries or action assignment during meetings.
  • Assigning work to absent colleagues, even if requested during the meeting.
  • Inferring deadlines or assigning owners based on job title.
  • External task-system integrations.
  • Automatically publishing unreviewed summaries.
  • Resolving identity using voice biometrics.

3. Users and primary workflow

Reviewer: Meeting organizer, or an explicitly designated attendee. Reviews and publishes the summary.

Attendee: Reads the summary, follows evidence links, and reports corrections.

Action owner: Receives published actions assigned to them.

Workflow:

  1. Meeting ends; Huddle completes transcript processing.
  2. Huddle reconciles actual attendance and generates a private draft.
  3. Reviewer checks decisions, unresolved ownership, and attendance exceptions.
  4. Reviewer edits and publishes.
  5. Huddle makes the summary available to authorized viewers and, where enabled, emails owners their approved actions.

If no eligible reviewer is available, the draft remains unpublished. A non-attending organizer may designate an attending reviewer but cannot resolve attendance exceptions themselves.

4. Attendance and identity requirements

Engineering must establish an attendance ledger separate from the calendar invite.

Each record contains:

  • Meeting ID and stable person ID, where resolved.
  • Join/leave intervals and participant-session IDs.
  • Identity source: authenticated join, verified guest identity, or shared-room roster.
  • Verification state and evidence.
  • Reviewer confirmation details, where required.

Calendar invitees are suggestions for identity matching, not proof of attendance. A directory search result is also not attendance evidence.

Assignment eligibility

An owner must map to a unique person with verified attendance in that meeting:

  • Authenticated individual join: Eligible when attendance telemetry records the person’s session.
  • Guest join: Eligible only after identity is resolved and verified; a display name alone is insufficient.
  • Shared meeting room: The room endpoint proves room participation, not individual presence. A participating reviewer must explicitly confirm each person physically present before they become eligible.
  • Uninvited attendee: Eligible through the same verification rules; invitation status is irrelevant.
  • Invited but absent person: Ineligible.

If telemetry is unavailable, identity is ambiguous, or attendance is disputed, block assignment. Keep the action unassigned with an explanation.

The owner picker must show only eligible attendees. Every write path—including generation, manual edits, APIs, and publication—must enforce eligibility server-side. Model confidence must never override this rule.

These controls cannot independently observe who is physically in a room. If trustworthy roster confirmation is unavailable, shared-room individuals remain ineligible rather than weakening the constraint.

5. Summary and extraction requirements

The draft has four sections:

  1. Overview: Two to four sentences describing the meeting’s purpose and outcome.
  2. Decisions: Clear statements of agreed outcomes.
  3. Action items: Task, owner or “Unassigned,” deadline if explicit, and status.
  4. Needs review: Uncertain decisions, ownership, identity, or attendance.

Each decision and action must include transcript-span references and recording timestamps. Preserve the original supporting text even when the reviewer edits the final wording.

Decisions

Capture explicit agreement or a clearly accepted conclusion. Do not convert brainstorming, recommendations, or unresolved debate into decisions. If discussion reverses an earlier decision, represent the final outcome and retain supporting evidence.

Actions

Extract a concrete deliverable or next step. Record a deadline only when explicitly stated; normalize relative dates using the meeting’s date and timezone.

Examples:

  • “I’ll send the revised deck tomorrow.” → Task with a candidate owner, subject to attribution review.
  • “Someone should update the deck.” → Unassigned action.
  • “Ask Priya to update it,” when Priya was absent → Unassigned action; note that an absent person was mentioned.
  • “We could revisit this next quarter.” → Not an action without a clear commitment.

Speaker labels provide evidence, not authority. Overlap, shared-room speech, ambiguous pronouns, and conflicting commitments must trigger ownership review. Candidate owners are draft suggestions, not assignments, until approved.

6. Review experience

The summary page opens in Draft state with a clear “Review before sharing” banner.

Design must provide:

  • Editable decision and action cards.
  • Evidence links opening the relevant transcript and recording segment.
  • Visible distinction between verified attendance and uncertain attribution.
  • “Unassigned” as a first-class owner option.
  • A dedicated queue for identity and attendance exceptions.
  • A publish preview showing recipients and notification settings.

For each assigned action, the reviewer must approve the owner. Bulk approval is permitted for straightforward items, but flagged attribution requires item-level review.

Attendance verification and ownership approval are separate controls: confirming that someone attended does not prove they accepted a task.

Publication is blocked if any assigned owner is ineligible. Unassigned actions and unresolved decision candidates do not block publication; uncertain decisions must be omitted from the final Decisions section or explicitly labeled unresolved.

7. Email policy and Sales request

Sales requests automatic emails immediately after each meeting. MVP will not send unreviewed action assignments. Given the known speaker errors, immediate delivery would spread incorrect ownership before correction.

Instead:

  • Notify the reviewer when the draft is ready.
  • Default owner emails to off; enable through workspace settings.
  • When enabled, automatically email approved owners after publication.
  • Send each owner only their actions, meeting title, deadlines, and an authorized summary link.
  • Never email a calendar invitee solely because they were invited.
  • Do not send to guests without a verified address and permitted access.

Email dispatch must recheck eligibility and publication version. Use idempotency keys to prevent duplicate sends. Published ownership changes generate a correction for the previous recipient and an updated notification for the new eligible owner.

Future immediate automation requires a separate launch decision and evidence that it preserves the attendance constraint.

8. Technical implementation

Use an asynchronous pipeline:

`Meeting ended → Transcript ready → Attendance reconciled → Draft generated → Reviewed → Published → Notifications`

Proposed records:

  • `Summary`: meeting ID, transcript version, state, generated content, revision.
  • `Decision`: text, evidence spans, review state.
  • `Action`: task, proposed owner, approved owner, deadline, evidence, review state.
  • `AttendanceRecord`: identity, evidence, verification state.
  • `AuditEvent`: actor, change, timestamp, revision.

Generation returns schema-validated structured data. Treat transcript content as untrusted input: instructions spoken during a meeting must not alter system rules, access controls, or recipient selection.

Owner writes and publication run through a centralized eligibility validator. Notification jobs check the current revision before sending. Attendance corrections invalidate affected assignments and cancel queued notifications. Maintain an audit trail rather than silently rewriting history.

9. Reliability, privacy, and access

Target draft availability: within five minutes of transcript completion for 95% of supported meetings.

Generation failures must leave recordings and transcripts usable, show retry status, and send no owner emails.

Drafts are visible only to reviewers. Published summaries inherit meeting-content access rules; attendance alone does not grant access. Recording evidence links must enforce existing permissions. Apply workspace retention and deletion policies to summaries, attendance evidence, and generated artifacts.

10. Success metrics and acceptance

Proposed launch gates

  • Zero assignments to ineligible attendees across automated invariant tests.
  • Zero emails to ineligible owners in the pilot; any confirmed incident pauses rollout.
  • At least 90% of published decisions require no substantive reviewer correction.
  • Median review time below three minutes for meetings under one hour.
  • At least 30% less self-reported note-taking time in the pilot.

Measure owner-correction rate separately from attendance violations: a wrong attending owner is still a quality failure.

Required acceptance cases

Test absent invitees, verified uninvited guests, shared rooms, overlapping speakers, duplicate names, failed identity resolution, missing telemetry, absent people named in requests, transcript reprocessing, attendance revoked before dispatch, and repeated notification retries.

In every uncertain-attendance case, expected behavior is an unassigned action and no owner email.

11. Rollout and dependencies

Start with internal meetings, then an opt-in pilot across varied company sizes and shared-room usage. Expand only after launch gates pass.

Before implementation, confirm availability of participant-session telemetry, guest identity verification, room-roster confirmation, reviewer permissions, and verified email addresses. If these dependencies are missing, ship summaries with unassigned actions for affected participants rather than substituting calendar membership as attendance proof.

Claims the judge couldn’t find in the brief

  • This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly12%
  • passAddresses the actual decision98%
  • passRespects explicit constraints68%
  • passIdentifies material uncertainty87%
  • failAvoids unsupported claims32%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs35%
  • partialProposes tests that could fail73%
Run
Run
#1
API response time
47 s
Submitted
29 Sept 2026

Needs more than a quick edit2 of 9 checks passed

It needs rework to require human review before auto-email and to replace unsupported attendance data assumptions. It cannot be used with only a quick edit.

Output

PRD: AI-Generated Meeting Summaries and Action Items

Audience: Engineering and design leads Product: Huddle Status: Draft

1. Summary

Huddle will generate a post-meeting summary that highlights decisions and action items, identifies owners when it can do so reliably, and can email confirmed action items to their owners automatically. This replaces the manual note-taking people do after calls.

The core safety requirement is non-negotiable: an action must never be assigned to someone who did not attend the meeting. Calendar invitations alone cannot establish attendance: some invitees do not join, and around 15% of meetings include someone who was not invited. Speaker labels also cannot establish identity on their own; they are wrong about 8% of the time, especially during interruptions or when several people share a meeting-room connection.

The product must therefore separate three questions: what was decided, what needs to be done, and who—if anyone—can be safely identified as the owner. When attendance or ownership is uncertain, Huddle must leave the action unassigned rather than guess.

2. Problem and goals

Today, attendees write their own notes after meetings. This is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.

Goals - Produce a concise, useful summary after a recorded meeting. - Clearly distinguish decisions from discussion, proposals, and unresolved questions. - Extract action items and assign them only to verified meeting attendees. - When enabled, email confirmed action items to their owners shortly after the meeting ends. - Let users correct mistakes and give feedback.

Non-goals for the first release - Replacing full meeting transcripts or recordings. - Creating tasks in external project-management tools. - Inferring deadlines, priorities, or owners that were not stated. - Automatically assigning work when the owner is unclear. - Resolving disagreements about whether a decision was made without review.

3. Users and primary workflow

Users: meeting attendees who need an accurate recap; organizers who want to review or share it; and action owners who need a clear follow-up.

  1. A recorded meeting ends. Huddle gathers the transcript, speaker labels, calendar context, and actual session attendance data.
  2. Huddle generates a summary with separate Decisions, Action items, and, where useful, Open questions sections.
  3. For each action, Huddle identifies the supporting transcript evidence and checks the proposed owner against the verified attendee roster.
  4. If automatic action-item email is enabled for the organization, Huddle emails only actions with a sufficiently clear owner whose attendance is verified.
  5. The organizer and attendees can view the summary, correct it, and report errors. Unassigned actions remain visible for follow-up but are not emailed as assigned work.

4. Functional requirements

4.1 Establishing who attended

  • Use Huddle’s meeting-session join/leave records as the attendance source of truth. The calendar invite is context, not proof that someone attended.
  • Include identifiable people who joined without a calendar invitation. Exclude invitees who did not join.
  • Represent an attendee using a verified Huddle account, authenticated guest identity, or another approved identity mechanism.
  • If a meeting-room device represents several people and Huddle cannot reliably map speakers to individuals, do not treat the device name or a guessed speaker identity as a person. The organizer may confirm identities or owners in the summary UI.
  • If reliable session attendance data is unavailable, fail closed: generate the summary, but do not assign or email action items.

4.2 Generating decisions and action items

  • Generate a short overview and the three optional sections: Decisions, Action items, and Open questions.
  • Include a decision only when the transcript supports an explicit agreement or choice. Label proposals or unresolved discussion as open questions, not decisions.
  • Phrase each action as a concrete task. Include a due date only if one was explicitly stated; otherwise show no due date.
  • Attach a brief transcript excerpt and timestamp to each decision and action so users can check the source.
  • Do not invent missing details. If the task is clear but no owner is, show it as Owner not confirmed. If the task itself is ambiguous, flag it for review rather than present it as certain.

4.3 Assigning owners safely

  • Every proposed owner must match a verified person in the actual attendee roster. Enforce this as a backend validation rule, not just a model instruction or UI check.
  • A calendar invitee who did not join must never be assigned. A person who joined without an invite may be assigned if their identity and attendance are verified.
  • Speaker labels may help identify who volunteered, but must not independently authorize an assignment. When overlap, room audio, or attribution uncertainty makes ownership unclear, leave the action unassigned.
  • Users may assign or change an owner in the UI only to a verified attendee. Record the change and its actor in the audit log.
  • Recheck the attendance invariant before saving an assignment and immediately before sending an email.

4.4 Automatic emails

  • Provide an organization-level setting to enable automatic action-item emails. Once enabled, send an email after generation to each verified owner with at least one confirmed action.
  • Do not send an assignment email for an unassigned action or to anyone not verified as having attended. Do not send the full transcript by default.
  • Include the meeting title and date, the recipient’s confirmed actions, any explicitly stated due dates, a link to the Huddle summary, and a way to report an incorrect assignment.
  • If the meeting has no verified owners, send no owner emails. The organizer may receive a link to the summary according to organization settings.
  • Target delivery within five minutes of the meeting ending at the 95th percentile. If generation or identity checks fail, do not send speculative emails; show the organizer that the summary or email is delayed or needs review.
  • Provide clear controls to turn automatic emails off. Email delivery status and failures must be visible to the organizer.

5. UX requirements

The summary should make confidence and uncertainty understandable without exposing raw model scores. Show an Owner confirmed state for verified assignments and Owner not confirmed when a person cannot be safely identified. Present source timestamps as links into the recording or transcript, subject to existing access permissions.

Organizers must be able to edit summary text, remove an incorrect item, or assign an unassigned item to a verified attendee. Edits must be saved with a clear indication of who made them. If an action email has already been sent, do not silently replace the original assignment: notify the organizer that the owner changed and offer a corrected email to the new verified owner. Never email the former owner again unless explicitly requested.

The UI should explain that the calendar list may include people who did not attend and that unidentified room participants require confirmation. Use clear empty states when there are no decisions or actions.

6. Data, privacy, and operational requirements

Use only meeting data already available to Huddle for recording, transcription, and attendance, subject to existing recording consent and organization policies. Apply existing access controls to summaries, transcript evidence, and recording links. Store generated summaries, edits, assignment records, email status, and relevant audit events under Huddle’s existing retention and deletion policies. Do not use action-item content to train models unless the organization’s existing data policy explicitly permits it.

Generation and email must be idempotent: retries must not create duplicate summary items or send duplicate emails. Log the evidence reference, proposed owner, roster-validation result, any human correction, and delivery outcome for troubleshooting and audit.

7. Acceptance criteria

The release is acceptable only if: 1. A calendar invitee who did not join is never assigned or emailed an action. 2. An identifiable, authenticated attendee who joined without an invite can be assigned when the transcript supports the assignment. 3. An unclear speaker in overlapping audio produces an unassigned action, not a guessed owner. 4. A shared-room connection with no reliable person mapping cannot result in an individual assignment without organizer confirmation. 5. Every emailed action has a verified attendee owner and a link to its supporting summary. 6. Email retries do not result in duplicate messages. 7. Users can correct an assignment, and the correction is restricted to verified attendees.

8. Success metrics and rollout

Track summary generation success, time to summary, owner-confirmation rate, action-item correction and removal rates, email delivery time, duplicate-email rate, and user feedback. The primary safety metric is the number of assignments or assignment emails to non-attendees; the target is zero. Also monitor missed or unassigned owners to understand the trade-off between safety and usefulness.

Roll out behind an organization-level feature flag. Start with internal meetings, then a small customer cohort. Review false assignments and user corrections before expanding. Do not broaden automatic-email availability until attendance validation, idempotency, and failure handling pass the acceptance criteria.

9. Dependency to resolve before implementation

Engineering must confirm that Huddle can provide reliable, person-level join/leave records, including authenticated guests and shared-room behavior. If that data is unavailable or insufficient for a meeting, the safe behavior is to generate the summary without assigning or emailing owners.

Claims the judge couldn’t find in the brief

  • Some invitees do not join meetings.
  • Huddle can provide reliable, person-level join/leave records as the attendance source of truth.
  • Today’s manual note-taking is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.
  • Huddle has existing meeting-session join/leave records and actual session attendance data.

Grades and run details

Decision model 72 · LLM judge 2 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly11%
  • passAddresses the actual decision90%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty86%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable78%
  • passSpecifies AI failure behaviour100%
  • partialSuccess settles trade-offs88%
  • partialProposes tests that could fail68%
Run
Run
#1
API response time
34 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 9 checks passed

All constraints are met, the reasoning is sound, and the PRD is detailed enough for engineering and design leads to start work with only minor editorial polish.

Output

Huddle AI meeting summaries

Product requirements for engineering and design leads

Purpose and product decision

Huddle will generate a reviewable record of meeting decisions and actions from its transcripts. The host checks the draft, resolves ownership and publishes it. Approved actions can then be emailed to their owners.

Huddle serves 3,000 companies whose users currently write notes themselves. Speaker labels are wrong about 8% of the time, and roughly 15% of meetings include an uninvited participant. Calendar membership and transcript attribution therefore cannot establish attendance or ownership.

The MVP will not fulfil Sales’ request for immediate, unreviewed owner emails. It will generate drafts automatically after transcription and send emails automatically after explicit publication. Faster delivery does not justify assigning work to absent people. All requirements and numerical targets below are proposed product decisions.

Scope and intended outcome

Primary users are meeting hosts reviewing records and attendees receiving actions. Success means less time producing dependable notes, with no assignments to non-attendees.

MVP includes English transcripts, decisions, action extraction, attendance verification, host review, publication and owner emails. Exclude live summaries, recurring reminders, task-system integrations and automatic reassignment. Use existing meeting access and retention controls.

An action has one verified owner or remains unassigned. Split genuinely separate responsibilities into separate actions. Do not infer deadlines, commitments or decisions from tentative discussion.

Attendance is a release dependency

Engineering must first confirm whether Huddle retains authenticated participant identities and join/leave events. The supplied calendar list is insufficient. If these records are unavailable, build attendance instrumentation before enabling assignments; summaries may still launch with unassigned actions.

An eligible owner must have a stable person identifier linked to an authenticated, human participant session with a recorded join event for this meeting occurrence. Calendar invitees who never joined are ineligible. Bots and room devices are ineligible. Deduplicate reconnects by identity. Someone who left early remains eligible, although a later discussion does not prove they accepted work.

Uninvited users are eligible when their authenticated join is recorded. Guest display names, shared-room speaker labels and email addresses typed by a reviewer are not identity evidence. Unverified guests and people represented only by a room device remain ineligible in MVP; explain how they can join individually with verified identity for future meetings.

This deliberately sacrifices assignment coverage. The invariant is enforceable against verified participation records, not proof of physical human presence behind an account. If literal physical presence is required, that remains an unresolved product constraint and assignments must not launch under a stronger guarantee.

Ownership and extraction rules

The model extracts structured candidates with supporting transcript spans. Each decision contains its outcome and evidence timestamps. Each action contains a task, supporting spans, optional explicit due date and a nullable candidate participant identifier. Preserve the original date wording alongside any normalised date, using the meeting timezone; ambiguous dates stay unset.

Only explicit commitments or clearly agreed assignments qualify as actions. Suggestions, questions and rejected proposals do not. Where later discussion reverses a decision, present the final supported outcome. If the resolution is unclear, flag it for review instead of declaring a decision.

A model candidate is internal metadata, never an assigned owner. Server validation rejects identifiers outside the eligible attendance set before anything is displayed. The review UI starts every owner field as “Unassigned”; reviewers choose an eligible attendee after checking evidence. Do not preselect names from speaker labels, including for “I’ll do it”.

Render ownership only from validated structured fields. Generated task and decision wording must not embed owner claims that bypass these controls. Unsupported person-specific claims are withheld for review. A request that absent Alex should do something can become an unassigned “Confirm responsibility for…” follow-up only when that follow-up was actually agreed; otherwise omit it as an action.

Review and publication experience

The meeting page shows “Preparing summary” after the call, then “Draft needs review”. The host is the default reviewer; existing authorised meeting editors may also review. If no reviewer is available, retain the draft without sending owner emails.

Display separate Decisions and Actions sections. Each item links to its transcript passage and recording timestamp. Offer edit, delete and add controls. Actions show task, owner and due date. An attendance-only owner picker explains why invitees or unresolved guests cannot be selected. Provide a visible “Needs owner” state rather than silently dropping useful work.

Reviewers must confirm each chosen owner and can publish with unresolved actions. The publish dialogue shows recipients and unresolved-action count, with an explicit “Publish and email owners” control. Actions without owners remain visible but produce no email. Keyboard navigation, labelled controls and screen-reader status announcements are required.

Published records show their revision and approval time. Subsequent edits create a new revision. Ownership changes require the same validation and explicit publication, then notify affected owners with a correction. Never silently replace a published record with regenerated content.

Engineering contract and failure behaviour

Use meeting occurrence IDs, not recurring calendar series IDs. Store transcript version, attendance evidence version, model/prompt version, summary revision, item IDs, evidence spans, owner ID, approval actor/time and delivery status. Keep internal candidates separate from approved owner fields.

Process states are waiting for transcript, generating, needs review, published and failed. Trigger generation only after transcript finalisation. Transcript changes invalidate unpublished approvals and create a fresh draft; published revisions remain intact until a replacement is approved.

Enforce owner eligibility server-side on draft writes, publication, exports and notification dispatch. Manual edits and direct API calls cannot bypass it. If attendance evidence is missing, contradictory or unavailable, block ownership and sending rather than falling back to calendar data. Record rejection reasons for support without exposing transcript content in routine logs.

Use transactional publication and an outbox. Delivery keys combine meeting, published revision and recipient so retries do not duplicate emails. Email only approved actions owned by that recipient, with a permission-checked meeting link. Resolve addresses from verified identities, never model output. Block delivery if permissions or identity validation fail; show the reviewer the failure. Retry transient failures with backoff for up to 24 hours, then offer manual retry.

Treat transcripts as untrusted data: spoken instructions cannot change permissions, choose recipients or trigger tools. Generation has no email capability. Preserve tenant isolation and apply meeting deletion and retention policies to summaries and queued notifications.

Acceptance scenarios

  1. Invited Pat never joins but is named in an action: Pat is absent from the picker; API assignment is rejected; no owner email is sent.
  1. Uninvited authenticated Lee joins: Lee is eligible and can receive an action after reviewer confirmation and publication.
  1. Several people join through one room device: no individual becomes eligible from its label; actions remain unassigned unless independent authenticated participation exists.
  1. Overlapping speech mislabels a commitment: no owner is preselected; the reviewer can inspect evidence and choose a verified attendee or leave it unassigned.
  1. No decisions or actions are supported: show explicit empty states rather than inventing content. Failed or incomplete transcription shows a retry state and sends nothing.
  1. A forged owner ID, missing attendance record or duplicate publish event cannot create an invalid assignment or duplicate email. Revoked access blocks queued delivery.

Evaluation and release

Build a consented, de-identified evaluation set of at least 200 meetings, deliberately including overlap, shared rooms, absent invitees, uninvited guests, duplicate names, reversals and ambiguous commitments. Human annotators label attendance, final decisions, actions, ownership and evidence; adjudicate disagreements independently. Report each difficult slice separately.

Release gates: zero non-attendee assignments or deliveries across end-to-end and adversarial tests; at least 95% precision and 85% recall for supported decisions and actions, measured separately. Zero observed violations is necessary but insufficient: eligibility enforcement must also pass API and notification-path tests. Model confidence alone never permits assignment.

Pilot with ten opt-in companies before staged rollout. Target draft availability within two minutes of final transcript for 95% of meetings up to 60 minutes, and a 50% reduction in median note-preparation time against pilot baseline. Track owner corrections, unassigned-action rate, publication rate and delivery failures. Any non-attendee assignment disables assignments and sending pending investigation. Expand only after two weeks without a critical violation and acceptable review effort; do not reduce abstention merely to improve coverage.

Grades and run details

Decision model 89 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision98%
  • passRespects explicit constraints62%
  • passIdentifies material uncertainty99%
  • partialAvoids unsupported claims21%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs92%
  • partialProposes tests that could fail58%
Run
Run
#1
Time to output
3.2 min
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongWrongRight
GPT-6.1 Sol · API

The output presents 'This creates duplicated effort, inconsistent records of decisions, and missed follow-through' as a current-situation fact, but it is not in the supplied context and does not follow by arithmetic.

GPT-6 Luna · API

It relies on unsupported current-situation claims such as reliable join/leave records and no-shows, which are not in the supplied context.

GPT-6 Astra · ChatGPT

All factual claims about the current situation come directly from the brief or context, with no invented numbers or false statements.

Respects explicit constraintsRightMixedRight
GPT-6.1 Sol · API

The output enforces the non-assignment constraint through an attendance ledger, eligibility checks, server-side validation, and reviewer approval; it also addresses the Sales request by deferring auto-email until after review.

GPT-6 Luna · API

It violates the hidden prohibition on auto-emailing action items without human review and likely exceeds the requested 1,500-word length.

GPT-6 Astra · ChatGPT

It respects the never‑assign‑to‑non‑attendees constraint through eligibility checks, host review, and server‑side enforcement; length, reader and form are all within the brief’s requirements.

Identifies material uncertaintyRightMixedRight
GPT-6.1 Sol · API

The output names specific unknowns such as shared-room physical attendance, missing telemetry, and ambiguous identity, and says how they would be resolved (reviewer confirmation, unassigned actions) and what would change the call (evidence for future automation).

GPT-6 Luna · API

It names speaker-label error and missing attendance data but does not specify attribution confidence thresholds or how those uncertainties would be resolved quantitatively.

GPT-6 Astra · ChatGPT

It identifies the unknown availability of authenticated join/leave records, bounds the impact, and states that if missing, assignments must be blocked until attendance instrumentation is built.

Avoids unsupported claimsWrongWrongRight
GPT-6.1 Sol · API

The output states 'This creates duplicated effort...' and 'immediate delivery would spread incorrect ownership' as established facts without labelling them as inferences or forecasts, although the supplied context only suggests these outcomes.

GPT-6 Luna · API

It presents unsupported claims about no-shows, join/leave records, and note-taking quality as established facts.

GPT-6 Astra · ChatGPT

Interpretations are clearly labelled as product decisions or derived from the supplied evidence; no confident claim goes beyond what the evidence supports.

Produces the required deliverableRightMixedRight
GPT-6.1 Sol · API

The output is a complete PRD aimed at engineering and design leads, covering decisions, actions, attribution, editing, and failure behaviour, and it appears to be within the requested 1,000-1,500 word range.

GPT-6 Luna · API

The PRD is complete in form but exceeds the requested length and depends on attendance data that is not confirmed to exist.

GPT-6 Astra · ChatGPT

The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.

Success settles trade-offsMixedWrongRight
GPT-6.1 Sol · API

The success metrics have numeric targets but no explicit time frame (e.g., pilot duration), and the output lacks a single clear trade-off rule stating which goal wins when precision and coverage conflict, though unassigned-over-misassignment is implied.

GPT-6 Luna · API

It gives a safety target and delivery target but lacks a time frame for success metrics and does not explicitly state which trade-off rule wins beyond vague monitoring.

GPT-6 Astra · ChatGPT

Success is defined with measurable targets (zero non‑attendee assignments, 95%/85% precision/recall, 50% reduction in note time) and an explicit trade‑off rule (sacrificing coverage to preserve correctness).

Proposes tests that could failWrongWrongRight
GPT-6.1 Sol · API

The launch gates and acceptance cases have numeric thresholds and some actions, but they lack measurement windows long enough to read out the effects, and not every gate specifies the action if the threshold is not met.

GPT-6 Luna · API

Acceptance criteria are mostly binary and lack numeric thresholds, measurement windows, or clear kill/pilot actions.

GPT-6 Astra · ChatGPT

Every acceptance scenario, release gate, and pilot has numeric thresholds, a clear measurement window, and a triggered action (e.g., disable assignments if any non‑attendee assignment occurs).

All got right 2

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The output commits early to 'Ship reviewed summaries first; defer immediate, unreviewed owner emails' and states that future immediate automation requires evidence that it preserves the attendance constraint.

GPT-6 Luna · API

It commits early to safe owner verification with optional auto-email and states conditions based on join/leave data availability.

GPT-6 Astra · ChatGPT

The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).

Specifies AI failure behaviourRightRightRight
GPT-6.1 Sol · API

The PRD explicitly specifies what happens when identity or attendance is uncertain (unassigned actions, no emails), when reviewers are unavailable (draft remains unpublished), and when generation fails (no owner emails, recordings usable).

GPT-6 Luna · API

It specifies fail-closed behavior, unassigned actions when uncertain, and no speculative emails.

GPT-6 Astra · ChatGPT

The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD