Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given4 items: Scenario, Known limits, Constraint, Request from Sales
ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer doesThe answer key the graders mark against

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score at 40)
  • Assigns actions to non-attendees
Case

Mid-level · v1.10 · synthetic · AI product, productivity

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit1 of 9 checks passed

A PM could not use it with only a quick edit because it relies on invented current systems and violates the required review-before-email constraint. The reasoning and evidence would need to be rechecked and the auto-dispatch design redone.

Output

Product Requirements Document (PRD)

Feature: Automated Meeting Summaries, Decisions, and Action Item Attribution Product: Huddle Video Conferencing Target Audience: Engineering Leads, Design Leads, Product Ops Status: Ready for Review Target Release: Q3

---

1. Executive Summary & Objective

Today, Huddle’s 3,000 corporate customers rely on manual note-taking. This results in lost context, untracked deliverables, and administrative drag.

This project introduces an AI-powered pipeline that automatically generates: 1. A concise meeting summary. 2. A structured list of key decisions. 3. Assigned action items mapped strictly to verified attendees.

Key Objectives

  • Reduce post-meeting admin time by 80%.
  • Achieve $\ge 90\%$ user acceptance rate on action item assignments (measured by lack of manual re-assignment or deletion).
  • Hard Rule Compliance: 0% assignment of action items to non-attendees.

---

2. Participant Verification & System Constraints

Our models must handle two structural data flaws: 1. Diarization/Labeling Error Rate: ~8% of transcript lines have inaccurate speaker tags (common during crosstalk or multi-person hardware rooms). 2. Ad-Hoc Joiners: ~15% of meetings feature participants who were not on the original calendar invite.

2.1 The "Eligible Assignee" Rule Engine (Non-Negotiable)

``` [Calendar Invite] [Client Telemetry Logs] \ / \ / v v +----------------------------------------+ | Participant Reconciliation Engine (PRE)| +----------------------------------------+ | v [Eligible Assignee Roster (EAR)] - Authenticated Users (User ID + Email) - Verified Guests (Session ID + Display Name) | v +----------------------------------------+ | LLM Extraction & Assignment Pipeline | +----------------------------------------+ | +--> Action assigned to User in EAR? --> ACCEPT | +--> Action assigned to External Party? | v Route to In-Room Proxy (or flag as Unassigned Dependency) ```

Constraint Definition: An action item can only be assigned to an entity present in the Eligible Assignee Roster (EAR) for that specific meeting session.

#### Step 1: Ingest Telemetry at Call Termination (`call.ended` event) Do not rely on the calendar invite as the source of truth for presence. Telemetry establishes physical attendance: * Calendar Attendees ($P{cal}$): Used only as an identity-matching fallback. * Connected Clients ($P{conn}$): Every participant with connection duration $> 60\text{ seconds}$ in the meeting telemetry logs. * Eligible Assignee Roster ($EAR$): Defined strictly as $EAR = P_{conn}$.

#### Step 2: Resolving the 15% Uninvited Attendees * Authenticated Joiners: If an uninvited joiner is logged into a Huddle account, resolve their `userid`, display name, and email directly into the $EAR$. * Guest Joiners (Unauthenticated): If a user joins via guest link without an account: * Capture their client display name and the session ID. * Tag them in the $EAR$ as `Guest: [Display Name]` with `userid = null`. * They are eligible for assignment in the UI, but will require an email address input if notifications are to be sent.

#### Step 3: Mitigating the 8% Diarization Noise Diarization errors mean the speaker label alone cannot be trusted for commitments (e.g., if Bob says "I'll take that," but the audio engine tags Alice). * The LLM prompt must cross-reference explicit verbal attribution ("Jane, can you take that?" $\rightarrow$ "Sure") against semantic context rather than relying exclusively on transcript speaker IDs. * When semantic confidence falls below threshold ($\tau < 0.75$), flag the action item as Unassigned rather than guessing.

#### Step 4: External Dependencies (Out-of-Room Third Parties) If a meeting participant says: "We need legal to review this, I'll have Sarah from Legal look at it," Sarah was not present and cannot be assigned the task. * Enforcement: The system assigns the task to the speaker (the in-room proxy) as: Assignee: Current Speaker (`[Speaker Name]`) Task: "Coordinate with Sarah (Legal) to review [X]." * If the speaker cannot be deduced with high confidence, set `assignee = UNASSIGNED` with a note: `External dependency: Sarah (Legal)`.

---

3. Product & Feature Specifications

3.1 LLM Processing Pipeline

Within 60 seconds of call termination, the processing service executes the following workflow:

``` [Call Ended] | v [Fetch Clean Transcript + Telemetry EAR] | v [LLM Context Construction] ├── System Prompt & Guardrails ├── Strict EAR List (Names + Emails + IDs) └── Diarized Transcript | v [Structured Output JSON Generation] ├── Executive Summary (Paragraph, <= 150 words) ├── Key Decisions (Bullet points) └── Action Items (Array: Task, Assignee ID, Due Date, Evidence Quote) | v [Deterministic Post-Processing Validator] ├── Cross-reference every Assignee ID against EAR └── Overwrite invalid assignments to "UNASSIGNED" | v [Summary Artifact Created & Persisted] ```

3.2 Action Item Post-Processing Validation Logic

Before writing to the database, run a deterministic verification function:

```python def validateactionitems(extractedactions, eligibleassigneeroster): validatedactions = [] for action in extractedactions: assigneeid = action.get("assigneeid") # Validation: Assignee must explicitly exist in the EAR if assigneeid and assigneeid in eligibleassigneeroster: action["status"] = "assigned" action["assignee"] = eligibleassigneeroster[assigneeid] else: # Fallback for hallucinated or external assignments action["status"] = "unassigned" action["assignee"] = None action["flag"] = "ASSIGNEENOTINROOM" validatedactions.append(action) return validated_actions ```

---

4. Sales Request Resolution: Auto-Email Notifications

The Conflict

Sales requested that action items be emailed automatically to assignees immediately when the meeting ends. However, because speaker labels fail ~8% of the time, immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

The Solution: "Grace Period Auto-Dispatch"

We satisfy Sales' need for immediate execution while protecting data integrity through a time-delayed trigger and host bypass mechanism.

``` [Meeting Ends] │ ├─► AI Generates Artifacts (Target: <60 seconds) │ ├─► Start 5-Minute Grace Period │ │ │ ├─► Host edits/approves? ──► Send instantly with updates │ │ │ └─► Timer expires? ────────► Auto-dispatch without host intervention │ └─► (Workspace Setting Override: "Instant Send" with 0-min delay available) ```

1. Host Review Window (Default): * A 5-minute countdown starts the moment the summary finishes processing. * The meeting host gets a browser/desktop push notification: "Your Huddle summary is ready. Auto-sending in 5:00 minutes. Click to review." * If the host makes edits or clicks "Send Now", the countdown ends and emails dispatch immediately. * If the host takes no action, the emails dispatch automatically at the 5-minute mark. 2. Sales/Workspace Configuration Flag: * Enterprise Admins can toggle the grace window per team (e.g., Sales team can set Grace Period to `0 minutes` for true instant dispatch; default for other workspaces is `5 minutes`). 3. Safety Fallback: Action items marked as `UNASSIGNED` are omitted from individual assignment emails and instead bundled exclusively into the primary meeting recap sent to the host and organizer.

---

5. User Experience & Design Specs

5.1 Post-Meeting In-App Summary Tab

Located inside the Huddle desktop and web apps under the meeting historical detail view.

``` +-----------------------------------------------------------------------+ | Meeting Summary: Product Architecture Sync [Share] [Edit]| | Today at 10:00 AM • 42 mins • 6 Attendees | +-----------------------------------------------------------------------+ | EXECUTIVE SUMMARY | | The team finalized the data migration strategy for Q3. Key risks | | around API rate limiting were addressed by introducing Redis caching.| +-----------------------------------------------------------------------+ | DECISIONS MADE | | • Chose Redis over Memcached for distributed cache. | | • Deferred mobile client updates to Sprint 44. | +-----------------------------------------------------------------------+ | ACTION ITEMS Auto-sending in [ 04:12 ] [x]| | [ + Add Action Item]| | [x] Task: Implement Redis cache layer | | Assignee: [ (Avatar) David Chen v ] Due: [ Friday v ] | | Source: "David: I can set up the cluster by end of week." | | | | [ ] Task: Verify Legal compliance on data retention | | Assignee: [ Unassigned v ] Due: [ Set Date v ] | | ! Flag: External dependency (Sarah / Legal was not present) | +-----------------------------------------------------------------------+ ```

5.2 Key UI Components

1. Assignee Dropdown Menu: * Filtered exclusively to individuals in the $EAR$. * Section 1: Authenticated Attendees (1-click select). * Section 2: Ad-Hoc / Guest Attendees (Displays `[Guest Name] - prompt for email`). * Option to select "Unassigned". * Explicitly disallows searching the global company directory within the "Assignee" field to prevent violating the non-attendee constraint. 2. Context Tooltip ("Source"): * Hovering over an action item shows the transcript snippet that triggered the assignment, allowing the host to quickly verify ambiguous attributions caused by diarization errors.

---

6. Email Dispatch Specs (Sales Notification System)

6.1 Individual Assignee Notification

  • To: Assignee email
  • Reply-To: Host email
  • Subject: `[Action Required] Action item from: {Meeting Title}`

``` Hi {Assignee First Name},

The following action item was assigned to you during {Meeting Title}:

Task: {Task Description} Due Date: {Due Date or "Not specified"} Context: "{Relevant transcript excerpt}"

View full meeting notes and recording: {Link to Huddle Meeting Record}

Did the AI get this wrong? Click here to reassign or remove: {One-Click Edit Link} ```

---

7. Telemetry, Edge Cases & Error Handling

Scenario / Edge CaseFailure Mode / RiskEngineering Resolution
:---:---:---
Multi-person room (Conference hardware)Single client IP/audio line represents 4 people; 1 is calendar-invited, 3 are walk-ins.PRE detects conference room profile. System assigns task to the Conference Room account by default, and sets an inline warning: "Assigned to hardware room. Click to reassign to an individual."
Cross-talk Diarization FlipsAudio assigns Alice's promise ("I'll ship it") to Bob.LLM applies verification: does speaker assignment align with prior conversation context? If confidence score $< 0.75$, fallback to `UNASSIGNED`.
Short/Accidental JoinsAttendee joined for 12 seconds, dropped, and was not present for decisions.PRE requires continuous session duration $\ge 60\text{ seconds}$ to enter the $EAR$.
Uninvited Guest Missing EmailUninvited guest gets an action item; auto-dispatch fails due to missing destination address.Action item is attributed to guest's display name. System flags: "Cannot send email: No address for [Display Name]." Host is prompted with an input box to provide an email.

---

8. Success Metrics

  1. Attribution Precision: $\ge 98\%$ of assigned action items correctly reflect real commitments made by that person (measured via sampling and low manual-edit rates).
  2. Zero-Hallucination Rate: $100\%$ compliance with the hard constraint: 0 automated emails dispatched to individuals outside the $EAR$.
  3. Dispatch Velocity: 95th percentile of summary and action item deliveries completed within $< 90\text{ seconds}$ post-meeting (or exactly at the expiration of the chosen grace period).
  4. Host Intervention Rate: $< 12\%$ of summaries require manual host corrections during the grace period window, indicating high trust in baseline accuracy.

Claims the judge couldn’t find in the brief

  • Manual note-taking results in lost context, untracked deliverables, and administrative drag.
  • Client telemetry logs exist and can establish physical attendance via connection duration greater than 60 seconds.
  • Authenticated users have resolvable user_id, display name, and email; guest joiners have session IDs and display names and may join via guest links.
  • Huddle has desktop and web apps with a meeting historical detail view.
  • There is a global company directory that could be searched for assignees.
  • Conference room accounts or profiles exist and can be assigned by default.
  • Immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

Grades and run details

Decision model 50 · LLM judge 1 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly67%
  • passAddresses the actual decision74%
  • failRespects explicit constraints63%
  • failIdentifies material uncertainty41%
  • failAvoids unsupported claims53%
  • passProduces the required deliverable66%
  • passSpecifies AI failure behaviour99%
  • passSuccess settles trade-offs31%
  • partialProposes tests that could fail67%
Run
Run
#1
API response time
35 s
Submitted
1 Oct 2026

Needs more than a quick edit2 of 9 checks passed

It needs rework to require human review before auto-email and to replace unsupported attendance data assumptions. It cannot be used with only a quick edit.

Output

PRD: AI-Generated Meeting Summaries and Action Items

Audience: Engineering and design leads Product: Huddle Status: Draft

1. Summary

Huddle will generate a post-meeting summary that highlights decisions and action items, identifies owners when it can do so reliably, and can email confirmed action items to their owners automatically. This replaces the manual note-taking people do after calls.

The core safety requirement is non-negotiable: an action must never be assigned to someone who did not attend the meeting. Calendar invitations alone cannot establish attendance: some invitees do not join, and around 15% of meetings include someone who was not invited. Speaker labels also cannot establish identity on their own; they are wrong about 8% of the time, especially during interruptions or when several people share a meeting-room connection.

The product must therefore separate three questions: what was decided, what needs to be done, and who—if anyone—can be safely identified as the owner. When attendance or ownership is uncertain, Huddle must leave the action unassigned rather than guess.

2. Problem and goals

Today, attendees write their own notes after meetings. This is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.

Goals - Produce a concise, useful summary after a recorded meeting. - Clearly distinguish decisions from discussion, proposals, and unresolved questions. - Extract action items and assign them only to verified meeting attendees. - When enabled, email confirmed action items to their owners shortly after the meeting ends. - Let users correct mistakes and give feedback.

Non-goals for the first release - Replacing full meeting transcripts or recordings. - Creating tasks in external project-management tools. - Inferring deadlines, priorities, or owners that were not stated. - Automatically assigning work when the owner is unclear. - Resolving disagreements about whether a decision was made without review.

3. Users and primary workflow

Users: meeting attendees who need an accurate recap; organizers who want to review or share it; and action owners who need a clear follow-up.

  1. A recorded meeting ends. Huddle gathers the transcript, speaker labels, calendar context, and actual session attendance data.
  2. Huddle generates a summary with separate Decisions, Action items, and, where useful, Open questions sections.
  3. For each action, Huddle identifies the supporting transcript evidence and checks the proposed owner against the verified attendee roster.
  4. If automatic action-item email is enabled for the organization, Huddle emails only actions with a sufficiently clear owner whose attendance is verified.
  5. The organizer and attendees can view the summary, correct it, and report errors. Unassigned actions remain visible for follow-up but are not emailed as assigned work.

4. Functional requirements

4.1 Establishing who attended

  • Use Huddle’s meeting-session join/leave records as the attendance source of truth. The calendar invite is context, not proof that someone attended.
  • Include identifiable people who joined without a calendar invitation. Exclude invitees who did not join.
  • Represent an attendee using a verified Huddle account, authenticated guest identity, or another approved identity mechanism.
  • If a meeting-room device represents several people and Huddle cannot reliably map speakers to individuals, do not treat the device name or a guessed speaker identity as a person. The organizer may confirm identities or owners in the summary UI.
  • If reliable session attendance data is unavailable, fail closed: generate the summary, but do not assign or email action items.

4.2 Generating decisions and action items

  • Generate a short overview and the three optional sections: Decisions, Action items, and Open questions.
  • Include a decision only when the transcript supports an explicit agreement or choice. Label proposals or unresolved discussion as open questions, not decisions.
  • Phrase each action as a concrete task. Include a due date only if one was explicitly stated; otherwise show no due date.
  • Attach a brief transcript excerpt and timestamp to each decision and action so users can check the source.
  • Do not invent missing details. If the task is clear but no owner is, show it as Owner not confirmed. If the task itself is ambiguous, flag it for review rather than present it as certain.

4.3 Assigning owners safely

  • Every proposed owner must match a verified person in the actual attendee roster. Enforce this as a backend validation rule, not just a model instruction or UI check.
  • A calendar invitee who did not join must never be assigned. A person who joined without an invite may be assigned if their identity and attendance are verified.
  • Speaker labels may help identify who volunteered, but must not independently authorize an assignment. When overlap, room audio, or attribution uncertainty makes ownership unclear, leave the action unassigned.
  • Users may assign or change an owner in the UI only to a verified attendee. Record the change and its actor in the audit log.
  • Recheck the attendance invariant before saving an assignment and immediately before sending an email.

4.4 Automatic emails

  • Provide an organization-level setting to enable automatic action-item emails. Once enabled, send an email after generation to each verified owner with at least one confirmed action.
  • Do not send an assignment email for an unassigned action or to anyone not verified as having attended. Do not send the full transcript by default.
  • Include the meeting title and date, the recipient’s confirmed actions, any explicitly stated due dates, a link to the Huddle summary, and a way to report an incorrect assignment.
  • If the meeting has no verified owners, send no owner emails. The organizer may receive a link to the summary according to organization settings.
  • Target delivery within five minutes of the meeting ending at the 95th percentile. If generation or identity checks fail, do not send speculative emails; show the organizer that the summary or email is delayed or needs review.
  • Provide clear controls to turn automatic emails off. Email delivery status and failures must be visible to the organizer.

5. UX requirements

The summary should make confidence and uncertainty understandable without exposing raw model scores. Show an Owner confirmed state for verified assignments and Owner not confirmed when a person cannot be safely identified. Present source timestamps as links into the recording or transcript, subject to existing access permissions.

Organizers must be able to edit summary text, remove an incorrect item, or assign an unassigned item to a verified attendee. Edits must be saved with a clear indication of who made them. If an action email has already been sent, do not silently replace the original assignment: notify the organizer that the owner changed and offer a corrected email to the new verified owner. Never email the former owner again unless explicitly requested.

The UI should explain that the calendar list may include people who did not attend and that unidentified room participants require confirmation. Use clear empty states when there are no decisions or actions.

6. Data, privacy, and operational requirements

Use only meeting data already available to Huddle for recording, transcription, and attendance, subject to existing recording consent and organization policies. Apply existing access controls to summaries, transcript evidence, and recording links. Store generated summaries, edits, assignment records, email status, and relevant audit events under Huddle’s existing retention and deletion policies. Do not use action-item content to train models unless the organization’s existing data policy explicitly permits it.

Generation and email must be idempotent: retries must not create duplicate summary items or send duplicate emails. Log the evidence reference, proposed owner, roster-validation result, any human correction, and delivery outcome for troubleshooting and audit.

7. Acceptance criteria

The release is acceptable only if: 1. A calendar invitee who did not join is never assigned or emailed an action. 2. An identifiable, authenticated attendee who joined without an invite can be assigned when the transcript supports the assignment. 3. An unclear speaker in overlapping audio produces an unassigned action, not a guessed owner. 4. A shared-room connection with no reliable person mapping cannot result in an individual assignment without organizer confirmation. 5. Every emailed action has a verified attendee owner and a link to its supporting summary. 6. Email retries do not result in duplicate messages. 7. Users can correct an assignment, and the correction is restricted to verified attendees.

8. Success metrics and rollout

Track summary generation success, time to summary, owner-confirmation rate, action-item correction and removal rates, email delivery time, duplicate-email rate, and user feedback. The primary safety metric is the number of assignments or assignment emails to non-attendees; the target is zero. Also monitor missed or unassigned owners to understand the trade-off between safety and usefulness.

Roll out behind an organization-level feature flag. Start with internal meetings, then a small customer cohort. Review false assignments and user corrections before expanding. Do not broaden automatic-email availability until attendance validation, idempotency, and failure handling pass the acceptance criteria.

9. Dependency to resolve before implementation

Engineering must confirm that Huddle can provide reliable, person-level join/leave records, including authenticated guests and shared-room behavior. If that data is unavailable or insufficient for a meeting, the safe behavior is to generate the summary without assigning or emailing owners.

Claims the judge couldn’t find in the brief

  • Some invitees do not join meetings.
  • Huddle can provide reliable, person-level join/leave records as the attendance source of truth.
  • Today’s manual note-taking is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.
  • Huddle has existing meeting-session join/leave records and actual session attendance data.

Grades and run details

Decision model 72 · LLM judge 2 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly11%
  • passAddresses the actual decision90%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty86%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable78%
  • passSpecifies AI failure behaviour100%
  • partialSuccess settles trade-offs88%
  • partialProposes tests that could fail68%
Run
Run
#1
API response time
34 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Addresses the actual decisionMixedRight
Gemini 3.8 Flash · API

It chooses a grace-period auto-dispatch answer but does not clearly state what result or condition would change that decision.

GPT-6 Luna · API

It commits early to safe owner verification with optional auto-email and states conditions based on join/leave data availability.

Respects explicit constraintsWrongMixed
Gemini 3.8 Flash · API

It violates the prohibition on auto-emailing without review by dispatching after five minutes even if no one reviews, and it likely exceeds the 1,000–1,500 word limit.

GPT-6 Luna · API

It violates the hidden prohibition on auto-emailing action items without human review and likely exceeds the requested 1,500-word length.

Identifies material uncertaintyWrongMixed
Gemini 3.8 Flash · API

It names known error rates but does not bound the material unknowns that could change the auto-email decision or specify what would change the call.

GPT-6 Luna · API

It names speaker-label error and missing attendance data but does not specify attribution confidence thresholds or how those uncertainties would be resolved quantitatively.

Success settles trade-offsMixedWrong
Gemini 3.8 Flash · API

It gives metric targets but no measurement time frame and no explicit rule for which goal wins when metrics conflict.

GPT-6 Luna · API

It gives a safety target and delivery target but lacks a time frame for success metrics and does not explicitly state which trade-off rule wins beyond vague monitoring.

All got wrong 3

Uses the supplied evidence correctlyWrongWrong
Gemini 3.8 Flash · API

It invents current systems and data such as telemetry logs, guest session IDs, desktop/web historical views, global directory, and conference room accounts that are not in the supplied context.

GPT-6 Luna · API

It relies on unsupported current-situation claims such as reliable join/leave records and no-shows, which are not in the supplied context.

Avoids unsupported claimsWrongWrong
Gemini 3.8 Flash · API

It presents several assumptions and forecasts as established facts, including telemetry availability, existing UI surfaces, and inevitable trust damage from immediate emails.

GPT-6 Luna · API

It presents unsupported claims about no-shows, join/leave records, and note-taking quality as established facts.

Proposes tests that could failWrongWrong
Gemini 3.8 Flash · API

Its gates and metrics lack measurement windows and clear actions for each pass/fail outcome.

GPT-6 Luna · API

Acceptance criteria are mostly binary and lack numeric thresholds, measurement windows, or clear kill/pilot actions.

All mixed 1

Produces the required deliverableMixedMixed
Gemini 3.8 Flash · API

Although it is a detailed PRD, it is too long for the requested range and contains major unsupported assumptions that engineering and design could not act on without re-checking.

GPT-6 Luna · API

The PRD is complete in form but exceeds the requested length and depends on attendance data that is not confirmed to exist.

All got right 1

Specifies AI failure behaviourRightRight
Gemini 3.8 Flash · API

It specifies low-confidence unassignment, deterministic validation, flags, editing, and missing-email handling.

GPT-6 Luna · API

It specifies fail-closed behavior, unassigned actions when uncertain, and no speculative emails.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD