Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given4 items: Scenario, Known limits, Constraint, Request from Sales
ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer doesThe answer key the graders mark against

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score at 40)
  • Assigns actions to non-attendees
Case

Mid-level · v1.10 · synthetic · AI product, productivity

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit1 of 9 checks passed

A PM could not use it with only a quick edit because it relies on invented current systems and violates the required review-before-email constraint. The reasoning and evidence would need to be rechecked and the auto-dispatch design redone.

Output

Product Requirements Document (PRD)

Feature: Automated Meeting Summaries, Decisions, and Action Item Attribution Product: Huddle Video Conferencing Target Audience: Engineering Leads, Design Leads, Product Ops Status: Ready for Review Target Release: Q3

---

1. Executive Summary & Objective

Today, Huddle’s 3,000 corporate customers rely on manual note-taking. This results in lost context, untracked deliverables, and administrative drag.

This project introduces an AI-powered pipeline that automatically generates: 1. A concise meeting summary. 2. A structured list of key decisions. 3. Assigned action items mapped strictly to verified attendees.

Key Objectives

  • Reduce post-meeting admin time by 80%.
  • Achieve $\ge 90\%$ user acceptance rate on action item assignments (measured by lack of manual re-assignment or deletion).
  • Hard Rule Compliance: 0% assignment of action items to non-attendees.

---

2. Participant Verification & System Constraints

Our models must handle two structural data flaws: 1. Diarization/Labeling Error Rate: ~8% of transcript lines have inaccurate speaker tags (common during crosstalk or multi-person hardware rooms). 2. Ad-Hoc Joiners: ~15% of meetings feature participants who were not on the original calendar invite.

2.1 The "Eligible Assignee" Rule Engine (Non-Negotiable)

``` [Calendar Invite] [Client Telemetry Logs] \ / \ / v v +----------------------------------------+ | Participant Reconciliation Engine (PRE)| +----------------------------------------+ | v [Eligible Assignee Roster (EAR)] - Authenticated Users (User ID + Email) - Verified Guests (Session ID + Display Name) | v +----------------------------------------+ | LLM Extraction & Assignment Pipeline | +----------------------------------------+ | +--> Action assigned to User in EAR? --> ACCEPT | +--> Action assigned to External Party? | v Route to In-Room Proxy (or flag as Unassigned Dependency) ```

Constraint Definition: An action item can only be assigned to an entity present in the Eligible Assignee Roster (EAR) for that specific meeting session.

#### Step 1: Ingest Telemetry at Call Termination (`call.ended` event) Do not rely on the calendar invite as the source of truth for presence. Telemetry establishes physical attendance: * Calendar Attendees ($P{cal}$): Used only as an identity-matching fallback. * Connected Clients ($P{conn}$): Every participant with connection duration $> 60\text{ seconds}$ in the meeting telemetry logs. * Eligible Assignee Roster ($EAR$): Defined strictly as $EAR = P_{conn}$.

#### Step 2: Resolving the 15% Uninvited Attendees * Authenticated Joiners: If an uninvited joiner is logged into a Huddle account, resolve their `userid`, display name, and email directly into the $EAR$. * Guest Joiners (Unauthenticated): If a user joins via guest link without an account: * Capture their client display name and the session ID. * Tag them in the $EAR$ as `Guest: [Display Name]` with `userid = null`. * They are eligible for assignment in the UI, but will require an email address input if notifications are to be sent.

#### Step 3: Mitigating the 8% Diarization Noise Diarization errors mean the speaker label alone cannot be trusted for commitments (e.g., if Bob says "I'll take that," but the audio engine tags Alice). * The LLM prompt must cross-reference explicit verbal attribution ("Jane, can you take that?" $\rightarrow$ "Sure") against semantic context rather than relying exclusively on transcript speaker IDs. * When semantic confidence falls below threshold ($\tau < 0.75$), flag the action item as Unassigned rather than guessing.

#### Step 4: External Dependencies (Out-of-Room Third Parties) If a meeting participant says: "We need legal to review this, I'll have Sarah from Legal look at it," Sarah was not present and cannot be assigned the task. * Enforcement: The system assigns the task to the speaker (the in-room proxy) as: Assignee: Current Speaker (`[Speaker Name]`) Task: "Coordinate with Sarah (Legal) to review [X]." * If the speaker cannot be deduced with high confidence, set `assignee = UNASSIGNED` with a note: `External dependency: Sarah (Legal)`.

---

3. Product & Feature Specifications

3.1 LLM Processing Pipeline

Within 60 seconds of call termination, the processing service executes the following workflow:

``` [Call Ended] | v [Fetch Clean Transcript + Telemetry EAR] | v [LLM Context Construction] ├── System Prompt & Guardrails ├── Strict EAR List (Names + Emails + IDs) └── Diarized Transcript | v [Structured Output JSON Generation] ├── Executive Summary (Paragraph, <= 150 words) ├── Key Decisions (Bullet points) └── Action Items (Array: Task, Assignee ID, Due Date, Evidence Quote) | v [Deterministic Post-Processing Validator] ├── Cross-reference every Assignee ID against EAR └── Overwrite invalid assignments to "UNASSIGNED" | v [Summary Artifact Created & Persisted] ```

3.2 Action Item Post-Processing Validation Logic

Before writing to the database, run a deterministic verification function:

```python def validateactionitems(extractedactions, eligibleassigneeroster): validatedactions = [] for action in extractedactions: assigneeid = action.get("assigneeid") # Validation: Assignee must explicitly exist in the EAR if assigneeid and assigneeid in eligibleassigneeroster: action["status"] = "assigned" action["assignee"] = eligibleassigneeroster[assigneeid] else: # Fallback for hallucinated or external assignments action["status"] = "unassigned" action["assignee"] = None action["flag"] = "ASSIGNEENOTINROOM" validatedactions.append(action) return validated_actions ```

---

4. Sales Request Resolution: Auto-Email Notifications

The Conflict

Sales requested that action items be emailed automatically to assignees immediately when the meeting ends. However, because speaker labels fail ~8% of the time, immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

The Solution: "Grace Period Auto-Dispatch"

We satisfy Sales' need for immediate execution while protecting data integrity through a time-delayed trigger and host bypass mechanism.

``` [Meeting Ends] │ ├─► AI Generates Artifacts (Target: <60 seconds) │ ├─► Start 5-Minute Grace Period │ │ │ ├─► Host edits/approves? ──► Send instantly with updates │ │ │ └─► Timer expires? ────────► Auto-dispatch without host intervention │ └─► (Workspace Setting Override: "Instant Send" with 0-min delay available) ```

1. Host Review Window (Default): * A 5-minute countdown starts the moment the summary finishes processing. * The meeting host gets a browser/desktop push notification: "Your Huddle summary is ready. Auto-sending in 5:00 minutes. Click to review." * If the host makes edits or clicks "Send Now", the countdown ends and emails dispatch immediately. * If the host takes no action, the emails dispatch automatically at the 5-minute mark. 2. Sales/Workspace Configuration Flag: * Enterprise Admins can toggle the grace window per team (e.g., Sales team can set Grace Period to `0 minutes` for true instant dispatch; default for other workspaces is `5 minutes`). 3. Safety Fallback: Action items marked as `UNASSIGNED` are omitted from individual assignment emails and instead bundled exclusively into the primary meeting recap sent to the host and organizer.

---

5. User Experience & Design Specs

5.1 Post-Meeting In-App Summary Tab

Located inside the Huddle desktop and web apps under the meeting historical detail view.

``` +-----------------------------------------------------------------------+ | Meeting Summary: Product Architecture Sync [Share] [Edit]| | Today at 10:00 AM • 42 mins • 6 Attendees | +-----------------------------------------------------------------------+ | EXECUTIVE SUMMARY | | The team finalized the data migration strategy for Q3. Key risks | | around API rate limiting were addressed by introducing Redis caching.| +-----------------------------------------------------------------------+ | DECISIONS MADE | | • Chose Redis over Memcached for distributed cache. | | • Deferred mobile client updates to Sprint 44. | +-----------------------------------------------------------------------+ | ACTION ITEMS Auto-sending in [ 04:12 ] [x]| | [ + Add Action Item]| | [x] Task: Implement Redis cache layer | | Assignee: [ (Avatar) David Chen v ] Due: [ Friday v ] | | Source: "David: I can set up the cluster by end of week." | | | | [ ] Task: Verify Legal compliance on data retention | | Assignee: [ Unassigned v ] Due: [ Set Date v ] | | ! Flag: External dependency (Sarah / Legal was not present) | +-----------------------------------------------------------------------+ ```

5.2 Key UI Components

1. Assignee Dropdown Menu: * Filtered exclusively to individuals in the $EAR$. * Section 1: Authenticated Attendees (1-click select). * Section 2: Ad-Hoc / Guest Attendees (Displays `[Guest Name] - prompt for email`). * Option to select "Unassigned". * Explicitly disallows searching the global company directory within the "Assignee" field to prevent violating the non-attendee constraint. 2. Context Tooltip ("Source"): * Hovering over an action item shows the transcript snippet that triggered the assignment, allowing the host to quickly verify ambiguous attributions caused by diarization errors.

---

6. Email Dispatch Specs (Sales Notification System)

6.1 Individual Assignee Notification

  • To: Assignee email
  • Reply-To: Host email
  • Subject: `[Action Required] Action item from: {Meeting Title}`

``` Hi {Assignee First Name},

The following action item was assigned to you during {Meeting Title}:

Task: {Task Description} Due Date: {Due Date or "Not specified"} Context: "{Relevant transcript excerpt}"

View full meeting notes and recording: {Link to Huddle Meeting Record}

Did the AI get this wrong? Click here to reassign or remove: {One-Click Edit Link} ```

---

7. Telemetry, Edge Cases & Error Handling

Scenario / Edge CaseFailure Mode / RiskEngineering Resolution
:---:---:---
Multi-person room (Conference hardware)Single client IP/audio line represents 4 people; 1 is calendar-invited, 3 are walk-ins.PRE detects conference room profile. System assigns task to the Conference Room account by default, and sets an inline warning: "Assigned to hardware room. Click to reassign to an individual."
Cross-talk Diarization FlipsAudio assigns Alice's promise ("I'll ship it") to Bob.LLM applies verification: does speaker assignment align with prior conversation context? If confidence score $< 0.75$, fallback to `UNASSIGNED`.
Short/Accidental JoinsAttendee joined for 12 seconds, dropped, and was not present for decisions.PRE requires continuous session duration $\ge 60\text{ seconds}$ to enter the $EAR$.
Uninvited Guest Missing EmailUninvited guest gets an action item; auto-dispatch fails due to missing destination address.Action item is attributed to guest's display name. System flags: "Cannot send email: No address for [Display Name]." Host is prompted with an input box to provide an email.

---

8. Success Metrics

  1. Attribution Precision: $\ge 98\%$ of assigned action items correctly reflect real commitments made by that person (measured via sampling and low manual-edit rates).
  2. Zero-Hallucination Rate: $100\%$ compliance with the hard constraint: 0 automated emails dispatched to individuals outside the $EAR$.
  3. Dispatch Velocity: 95th percentile of summary and action item deliveries completed within $< 90\text{ seconds}$ post-meeting (or exactly at the expiration of the chosen grace period).
  4. Host Intervention Rate: $< 12\%$ of summaries require manual host corrections during the grace period window, indicating high trust in baseline accuracy.

Claims the judge couldn’t find in the brief

  • Manual note-taking results in lost context, untracked deliverables, and administrative drag.
  • Client telemetry logs exist and can establish physical attendance via connection duration greater than 60 seconds.
  • Authenticated users have resolvable user_id, display name, and email; guest joiners have session IDs and display names and may join via guest links.
  • Huddle has desktop and web apps with a meeting historical detail view.
  • There is a global company directory that could be searched for assignees.
  • Conference room accounts or profiles exist and can be assigned by default.
  • Immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

Grades and run details

Decision model 50 · LLM judge 1 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly67%
  • passAddresses the actual decision74%
  • failRespects explicit constraints63%
  • failIdentifies material uncertainty41%
  • failAvoids unsupported claims53%
  • passProduces the required deliverable66%
  • passSpecifies AI failure behaviour99%
  • passSuccess settles trade-offs31%
  • partialProposes tests that could fail67%
Run
Run
#1
API response time
35 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 9 checks passed

All constraints are met, the reasoning is sound, and the PRD is detailed enough for engineering and design leads to start work with only minor editorial polish.

Output

Huddle AI meeting summaries

Product requirements for engineering and design leads

Purpose and product decision

Huddle will generate a reviewable record of meeting decisions and actions from its transcripts. The host checks the draft, resolves ownership and publishes it. Approved actions can then be emailed to their owners.

Huddle serves 3,000 companies whose users currently write notes themselves. Speaker labels are wrong about 8% of the time, and roughly 15% of meetings include an uninvited participant. Calendar membership and transcript attribution therefore cannot establish attendance or ownership.

The MVP will not fulfil Sales’ request for immediate, unreviewed owner emails. It will generate drafts automatically after transcription and send emails automatically after explicit publication. Faster delivery does not justify assigning work to absent people. All requirements and numerical targets below are proposed product decisions.

Scope and intended outcome

Primary users are meeting hosts reviewing records and attendees receiving actions. Success means less time producing dependable notes, with no assignments to non-attendees.

MVP includes English transcripts, decisions, action extraction, attendance verification, host review, publication and owner emails. Exclude live summaries, recurring reminders, task-system integrations and automatic reassignment. Use existing meeting access and retention controls.

An action has one verified owner or remains unassigned. Split genuinely separate responsibilities into separate actions. Do not infer deadlines, commitments or decisions from tentative discussion.

Attendance is a release dependency

Engineering must first confirm whether Huddle retains authenticated participant identities and join/leave events. The supplied calendar list is insufficient. If these records are unavailable, build attendance instrumentation before enabling assignments; summaries may still launch with unassigned actions.

An eligible owner must have a stable person identifier linked to an authenticated, human participant session with a recorded join event for this meeting occurrence. Calendar invitees who never joined are ineligible. Bots and room devices are ineligible. Deduplicate reconnects by identity. Someone who left early remains eligible, although a later discussion does not prove they accepted work.

Uninvited users are eligible when their authenticated join is recorded. Guest display names, shared-room speaker labels and email addresses typed by a reviewer are not identity evidence. Unverified guests and people represented only by a room device remain ineligible in MVP; explain how they can join individually with verified identity for future meetings.

This deliberately sacrifices assignment coverage. The invariant is enforceable against verified participation records, not proof of physical human presence behind an account. If literal physical presence is required, that remains an unresolved product constraint and assignments must not launch under a stronger guarantee.

Ownership and extraction rules

The model extracts structured candidates with supporting transcript spans. Each decision contains its outcome and evidence timestamps. Each action contains a task, supporting spans, optional explicit due date and a nullable candidate participant identifier. Preserve the original date wording alongside any normalised date, using the meeting timezone; ambiguous dates stay unset.

Only explicit commitments or clearly agreed assignments qualify as actions. Suggestions, questions and rejected proposals do not. Where later discussion reverses a decision, present the final supported outcome. If the resolution is unclear, flag it for review instead of declaring a decision.

A model candidate is internal metadata, never an assigned owner. Server validation rejects identifiers outside the eligible attendance set before anything is displayed. The review UI starts every owner field as “Unassigned”; reviewers choose an eligible attendee after checking evidence. Do not preselect names from speaker labels, including for “I’ll do it”.

Render ownership only from validated structured fields. Generated task and decision wording must not embed owner claims that bypass these controls. Unsupported person-specific claims are withheld for review. A request that absent Alex should do something can become an unassigned “Confirm responsibility for…” follow-up only when that follow-up was actually agreed; otherwise omit it as an action.

Review and publication experience

The meeting page shows “Preparing summary” after the call, then “Draft needs review”. The host is the default reviewer; existing authorised meeting editors may also review. If no reviewer is available, retain the draft without sending owner emails.

Display separate Decisions and Actions sections. Each item links to its transcript passage and recording timestamp. Offer edit, delete and add controls. Actions show task, owner and due date. An attendance-only owner picker explains why invitees or unresolved guests cannot be selected. Provide a visible “Needs owner” state rather than silently dropping useful work.

Reviewers must confirm each chosen owner and can publish with unresolved actions. The publish dialogue shows recipients and unresolved-action count, with an explicit “Publish and email owners” control. Actions without owners remain visible but produce no email. Keyboard navigation, labelled controls and screen-reader status announcements are required.

Published records show their revision and approval time. Subsequent edits create a new revision. Ownership changes require the same validation and explicit publication, then notify affected owners with a correction. Never silently replace a published record with regenerated content.

Engineering contract and failure behaviour

Use meeting occurrence IDs, not recurring calendar series IDs. Store transcript version, attendance evidence version, model/prompt version, summary revision, item IDs, evidence spans, owner ID, approval actor/time and delivery status. Keep internal candidates separate from approved owner fields.

Process states are waiting for transcript, generating, needs review, published and failed. Trigger generation only after transcript finalisation. Transcript changes invalidate unpublished approvals and create a fresh draft; published revisions remain intact until a replacement is approved.

Enforce owner eligibility server-side on draft writes, publication, exports and notification dispatch. Manual edits and direct API calls cannot bypass it. If attendance evidence is missing, contradictory or unavailable, block ownership and sending rather than falling back to calendar data. Record rejection reasons for support without exposing transcript content in routine logs.

Use transactional publication and an outbox. Delivery keys combine meeting, published revision and recipient so retries do not duplicate emails. Email only approved actions owned by that recipient, with a permission-checked meeting link. Resolve addresses from verified identities, never model output. Block delivery if permissions or identity validation fail; show the reviewer the failure. Retry transient failures with backoff for up to 24 hours, then offer manual retry.

Treat transcripts as untrusted data: spoken instructions cannot change permissions, choose recipients or trigger tools. Generation has no email capability. Preserve tenant isolation and apply meeting deletion and retention policies to summaries and queued notifications.

Acceptance scenarios

  1. Invited Pat never joins but is named in an action: Pat is absent from the picker; API assignment is rejected; no owner email is sent.
  1. Uninvited authenticated Lee joins: Lee is eligible and can receive an action after reviewer confirmation and publication.
  1. Several people join through one room device: no individual becomes eligible from its label; actions remain unassigned unless independent authenticated participation exists.
  1. Overlapping speech mislabels a commitment: no owner is preselected; the reviewer can inspect evidence and choose a verified attendee or leave it unassigned.
  1. No decisions or actions are supported: show explicit empty states rather than inventing content. Failed or incomplete transcription shows a retry state and sends nothing.
  1. A forged owner ID, missing attendance record or duplicate publish event cannot create an invalid assignment or duplicate email. Revoked access blocks queued delivery.

Evaluation and release

Build a consented, de-identified evaluation set of at least 200 meetings, deliberately including overlap, shared rooms, absent invitees, uninvited guests, duplicate names, reversals and ambiguous commitments. Human annotators label attendance, final decisions, actions, ownership and evidence; adjudicate disagreements independently. Report each difficult slice separately.

Release gates: zero non-attendee assignments or deliveries across end-to-end and adversarial tests; at least 95% precision and 85% recall for supported decisions and actions, measured separately. Zero observed violations is necessary but insufficient: eligibility enforcement must also pass API and notification-path tests. Model confidence alone never permits assignment.

Pilot with ten opt-in companies before staged rollout. Target draft availability within two minutes of final transcript for 95% of meetings up to 60 minutes, and a 50% reduction in median note-preparation time against pilot baseline. Track owner corrections, unassigned-action rate, publication rate and delivery failures. Any non-attendee assignment disables assignments and sending pending investigation. Expand only after two weeks without a critical violation and acceptable review effort; do not reduce abstention merely to improve coverage.

Grades and run details

Decision model 89 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision98%
  • passRespects explicit constraints62%
  • passIdentifies material uncertainty99%
  • partialAvoids unsupported claims21%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs92%
  • partialProposes tests that could fail58%
Run
Run
#1
Time to output
3.2 min
Submitted
24 Sept 2026

Needs more than a quick edit5 of 9 checks passed

It includes an unsupported claim about note quality that a product manager would need to verify, and it omits a time frame for its success metrics, requiring more than a quick edit.

Output

PRD: AI Meeting Summaries — Decisions & Action Items

Sep 24, 2026 · @Martin Slaney

Summary

We will generate a summary after every recorded Huddle meeting that lists the decisions made and the action items, each with an owner who was verifiably in the meeting. The organizer reviews the summary and sends it in one click; nothing is emailed without that step in v1.

Problem. Today attendees write their own notes after the call. Notes are inconsistent, often missing, and action items lose their owners within a day. We already have what we need to fix this: a recording, a transcript with speaker labels, and the calendar attendee list.

Why this is not trivial. Two of those inputs are unreliable. Speaker labels are wrong about 8% of the time, mostly during crosstalk and when several people share a meeting-room device. About 15% of meetings include someone who joined without a calendar invite. A naive build would misassign owners often enough to lose trust in the first week.

Goals and metrics

The goal is summaries people send without rewriting, with zero actions assigned to non-attendees.

Goals

  • Replace manual post-meeting notes for recorded meetings.
  • Every action item has a correct owner or is explicitly marked unassigned.
  • Get action items to owners within minutes of the meeting ending.

Non-goals (v1)

  • Pushing tasks into Jira, Asana or Linear.
  • Summaries for unrecorded meetings or live, in-call summaries.
  • Improving diarization itself. We design around the 8% error rate, not fix it.

Success metrics

MetricTargetType
Actions assigned to someone not in the meeting0Hard constraint, blocks release
Owner precision on items sent with an owner≥ 97%Quality
Summaries sent within 30 min of meeting end≥ 60% of recorded meetingsAdoption
Summaries sent with no edits≥ 50%Quality
Owner reassignments per sent summary≤ 0.3Quality

Owner precision is measured from organizer edits plus a weekly labelled sample of 200 meetings.

Users and core flow

The organizer owns the summary; attendees receive it.

  • Organizer: reviews, fixes owners, sends. Usually the calendar invite's creator; can hand off to any attendee.
  • Action owner: receives their items by email and in the app.
  • Other attendees: receive the full summary once sent.

```mermaid flowchart LR A[Meeting ends] --> B[Transcript +<br/>join log ready] B --> C[Build verified<br/>roster] C --> D[Extract decisions<br/>and actions] D --> E[Roster filter +<br/>confidence check] E --> F[Organizer review] F --> G[Send to owners<br/>and attendees] ```

The roster filter runs as deterministic code after the model, never inside the prompt alone.2 Target: summary ready for review within 5 minutes of meeting end for meetings up to 90 minutes.

Requirements

An owner can only be chosen from the verified roster: people who actually joined, not people who were invited.

Verified roster

1. R1. Build the roster from the join log, not the calendar invite.1 Invited people who never joined are excluded. 2. R2. Each roster entry has an identity tier: - Verified: joined signed in to a Huddle account, or by a dial-in matched to a known number. - Guest: joined with a display name only (most of the 15% uninvited joiners). - Room: a meeting-room device; people behind it are unknown individuals. 3. R3. Match uninvited joiners to workspace accounts by signed-in account only. Never match on display name alone.

Summary content

  1. R4. Decisions: a short statement of each decision, with a transcript timestamp link. No owner field.
  2. R5. Action items: task, owner (or "Unassigned"), due date if said aloud, timestamp link, and confidence.
  3. R6. Every item links to its evidence in the transcript so reviewers can check it in one click.

Assignment rules

  1. R7. The model returns a proposed owner per item. Code then checks it against the roster. Any owner not on the roster is replaced with "Unassigned", and the named person is shown as a mention ("mentions: Dave, not in meeting").
  2. R8. Owner comes from content first: a named request ("Priya, can you…") beats the speaker label of "I'll do it".
  3. R9. Mark an owner low confidence when the evidence depends on a speaker label in an overlapping segment, a Room entry, or a Guest entry.
  4. R10. Low-confidence items show as "Suggested: Priya" and need organizer confirmation before any email goes out.
  5. R11. Room and Guest entries cannot receive email. The organizer can reassign to a verified person or leave it unassigned.
  6. R12. Log every model proposal, filter result and organizer edit. This is our precision measurement.

Handling the known limits

Both limits hit ownership, not decisions, so we contain them at the owner step.

LimitWhat goes wrongHow we handle it
Speaker labels wrong \~8% (crosstalk)"I'll take that" credited to the wrong personContent-first owner rule (R8); overlap segments flagged low confidence (R9)
Speaker labels wrong \~8% (shared room)One label covers several peopleRoom tier; items resolved to a room need a named person or stay unassigned (R11)
\~15% of meetings have uninvited joinersReal participant missing from roster, or matched to the wrong accountRoster from join log (R1); account match only (R3); Guests can own items but get no email
Invited but absentModel assigns to a name from the inviteExcluded from roster; hard filter (R7)

The error compounds per meeting. If each item's owner is wrong 8% of the time, a meeting with 5 action items has a 34% chance of at least one wrong owner (1 − 0.92⁵). That is why confirmation sits before email, not after.

Sales request: auto-email on meeting end

Decision: not in v1 as requested. We ship one-click organizer send now, and gated auto-send in phase 2 once measured precision earns it.

Why not now. An email to the wrong owner is the most visible failure this feature can have. With the error rate above, roughly a third of 5-item meetings would email someone a task that isn't theirs. With external guests, it may land outside the customer's company. Early trust is hard to win back.

What Sales gets instead.

  • v1: the organizer gets a "Summary ready" notification within 5 minutes and sends with one click. Goal: sent within 30 minutes (see metrics).
  • Phase 2: a workspace setting to auto-send at meeting end, but only items that are high confidence and owned by a Verified attendee. Everything else still waits for the organizer. It ships only when owner precision on high-confidence items is ≥ 97% for 4 consecutive weeks.
  • Auto-send never emails Guest or Room entries, or anyone outside the customer's domain, unless an admin enables it.

Design notes

The review screen must make checking owners faster than writing notes.

  • Decisions first, then action items grouped by owner, then Unassigned.
  • Low-confidence owners look visibly different ("Suggested") and are editable inline from a roster picker, never free text.
  • Each item plays its transcript clip on click.
  • The Send button states who gets email: "Send to 4 owners and 7 attendees".
  • Empty state: if no decisions or actions are found, say so plainly rather than inventing any.

Rollout

PhaseScopeExit criteria
InternalHuddle's own meetingsZero non-attendee assignments; precision logged
Beta50 opted-in companies≥ 97% owner precision on sent items; ≥ 40% sent within 30 min
GAAll 3,000 companies, organizer sendBeta targets held for 4 weeks
Phase 2Gated auto-send settingSee Sales section

Open questions

  • [ ] Is the join log's signed-in identity reliable enough to call "Verified"? Engineering to confirm.
  • [ ] Who organizes when the invite creator didn't attend? Proposed: first Verified attendee from the invite.
  • [ ] Do customers need to disable summaries per meeting for sensitive calls? Legal and Sales to confirm.
  • [ ] Retention: how long do we keep drafts and the evidence log?

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    Source: Scenariothe attendee list comes from the calendar invite

    What we’d changeTreat the join log as a dependency to confirm, not an input that exists today. Say what proves attendance, and don't promise zero non-attendee assignments until that's settled.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: ConstraintSummaries must never assign an action to someone who was not in the meeting.

    What we’d changeApply the roster check to organiser edits and the final send as well, not only to the model's output.

Claims the judge couldn’t find in the brief

  • Notes are inconsistent, often missing, and action items lose their owners within a day.

Grades and run details

Decision model 72 · LLM judge 5 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly43%
  • passAddresses the actual decision98%
  • passRespects explicit constraints68%
  • passIdentifies material uncertainty89%
  • failAvoids unsupported claims38%
  • passProduces the required deliverable78%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs94%
  • partialProposes tests that could fail61%
Artefacts
Run
Run
#1
Time to output
80 s
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyWrongRightWrong
Gemini 3.8 Flash · API

It invents current systems and data such as telemetry logs, guest session IDs, desktop/web historical views, global directory, and conference room accounts that are not in the supplied context.

GPT-6 Astra · ChatGPT

All factual claims about the current situation come directly from the brief or context, with no invented numbers or false statements.

Opus 5.5 · Claude

The output claims that today's notes are inconsistent, often missing, and action items lose owners within a day; this is not in the supplied context and invents current performance.

Addresses the actual decisionMixedRightRight
Gemini 3.8 Flash · API

It chooses a grace-period auto-dispatch answer but does not clearly state what result or condition would change that decision.

GPT-6 Astra · ChatGPT

The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).

Opus 5.5 · Claude

It commits early to organizer-reviewed summaries with no auto-email in v1, says what would change that (precision data), and frames the choice for engineering and design leads.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It violates the prohibition on auto-emailing without review by dispatching after five minutes even if no one reviews, and it likely exceeds the 1,000–1,500 word limit.

GPT-6 Astra · ChatGPT

It respects the never‑assign‑to‑non‑attendees constraint through eligibility checks, host review, and server‑side enforcement; length, reader and form are all within the brief’s requirements.

Opus 5.5 · Claude

The PRD enforces the zero-non-attendee constraint with a verified roster and hard filter, and the output is within the 1,000–1,500 word aim.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It names known error rates but does not bound the material unknowns that could change the auto-email decision or specify what would change the call.

GPT-6 Astra · ChatGPT

It identifies the unknown availability of authenticated join/leave records, bounds the impact, and states that if missing, assignments must be blocked until attendance instrumentation is built.

Opus 5.5 · Claude

It names specific unknowns (join log reliability, organizer when creator absent, per-meeting disable, retention) and how they would be resolved.

Avoids unsupported claimsWrongRightWrong
Gemini 3.8 Flash · API

It presents several assumptions and forecasts as established facts, including telemetry availability, existing UI surfaces, and inevitable trust damage from immediate emails.

GPT-6 Astra · ChatGPT

Interpretations are clearly labelled as product decisions or derived from the supplied evidence; no confident claim goes beyond what the evidence supports.

Opus 5.5 · Claude

The statement about notes being inconsistent and action items losing owners is presented as fact without the supplied evidence supporting it.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

Although it is a detailed PRD, it is too long for the requested range and contains major unsupported assumptions that engineering and design could not act on without re-checking.

GPT-6 Astra · ChatGPT

The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.

Opus 5.5 · Claude

The output is a PRD that gives requirements, flow, design notes, and rollout; engineering and design leads could start work from it with light edits.

Success settles trade-offsMixedRightMixed
Gemini 3.8 Flash · API

It gives metric targets but no measurement time frame and no explicit rule for which goal wins when metrics conflict.

GPT-6 Astra · ChatGPT

Success is defined with measurable targets (zero non‑attendee assignments, 95%/85% precision/recall, 50% reduction in note time) and an explicit trade‑off rule (sacrificing coverage to preserve correctness).

Opus 5.5 · Claude

Success metrics lack a time frame, and no explicit trade-off rule is stated beyond the hard constraint, so the PRD does not define which goal wins when two conflict.

Proposes tests that could failWrongRightWrong
Gemini 3.8 Flash · API

Its gates and metrics lack measurement windows and clear actions for each pass/fail outcome.

GPT-6 Astra · ChatGPT

Every acceptance scenario, release gate, and pilot has numeric thresholds, a clear measurement window, and a triggered action (e.g., disable assignments if any non‑attendee assignment occurs).

Opus 5.5 · Claude

Internal and beta exit criteria lack a measurement window (beta has no duration, internal only 'zero non‑attendee assignments' without a window), and not all phases state the action each result triggers.

All got right 1

Specifies AI failure behaviourRightRightRight
Gemini 3.8 Flash · API

It specifies low-confidence unassignment, deterministic validation, flags, editing, and missing-email handling.

GPT-6 Astra · ChatGPT

The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.

Opus 5.5 · Claude

It specifies low‑confidence flags, unassigned fallback, and organizer confirmation before email, covering AI uncertainty and errors.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD