Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given4 items: Scenario, Known limits, Constraint, Request from Sales
ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer doesThe answer key the graders mark against

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score at 40)
  • Assigns actions to non-attendees
Case

Mid-level · v1.10 · synthetic · AI product, productivity

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit1 of 9 checks passed

A PM could not use it with only a quick edit because it relies on invented current systems and violates the required review-before-email constraint. The reasoning and evidence would need to be rechecked and the auto-dispatch design redone.

Output

Product Requirements Document (PRD)

Feature: Automated Meeting Summaries, Decisions, and Action Item Attribution Product: Huddle Video Conferencing Target Audience: Engineering Leads, Design Leads, Product Ops Status: Ready for Review Target Release: Q3

---

1. Executive Summary & Objective

Today, Huddle’s 3,000 corporate customers rely on manual note-taking. This results in lost context, untracked deliverables, and administrative drag.

This project introduces an AI-powered pipeline that automatically generates: 1. A concise meeting summary. 2. A structured list of key decisions. 3. Assigned action items mapped strictly to verified attendees.

Key Objectives

  • Reduce post-meeting admin time by 80%.
  • Achieve $\ge 90\%$ user acceptance rate on action item assignments (measured by lack of manual re-assignment or deletion).
  • Hard Rule Compliance: 0% assignment of action items to non-attendees.

---

2. Participant Verification & System Constraints

Our models must handle two structural data flaws: 1. Diarization/Labeling Error Rate: ~8% of transcript lines have inaccurate speaker tags (common during crosstalk or multi-person hardware rooms). 2. Ad-Hoc Joiners: ~15% of meetings feature participants who were not on the original calendar invite.

2.1 The "Eligible Assignee" Rule Engine (Non-Negotiable)

``` [Calendar Invite] [Client Telemetry Logs] \ / \ / v v +----------------------------------------+ | Participant Reconciliation Engine (PRE)| +----------------------------------------+ | v [Eligible Assignee Roster (EAR)] - Authenticated Users (User ID + Email) - Verified Guests (Session ID + Display Name) | v +----------------------------------------+ | LLM Extraction & Assignment Pipeline | +----------------------------------------+ | +--> Action assigned to User in EAR? --> ACCEPT | +--> Action assigned to External Party? | v Route to In-Room Proxy (or flag as Unassigned Dependency) ```

Constraint Definition: An action item can only be assigned to an entity present in the Eligible Assignee Roster (EAR) for that specific meeting session.

#### Step 1: Ingest Telemetry at Call Termination (`call.ended` event) Do not rely on the calendar invite as the source of truth for presence. Telemetry establishes physical attendance: * Calendar Attendees ($P{cal}$): Used only as an identity-matching fallback. * Connected Clients ($P{conn}$): Every participant with connection duration $> 60\text{ seconds}$ in the meeting telemetry logs. * Eligible Assignee Roster ($EAR$): Defined strictly as $EAR = P_{conn}$.

#### Step 2: Resolving the 15% Uninvited Attendees * Authenticated Joiners: If an uninvited joiner is logged into a Huddle account, resolve their `userid`, display name, and email directly into the $EAR$. * Guest Joiners (Unauthenticated): If a user joins via guest link without an account: * Capture their client display name and the session ID. * Tag them in the $EAR$ as `Guest: [Display Name]` with `userid = null`. * They are eligible for assignment in the UI, but will require an email address input if notifications are to be sent.

#### Step 3: Mitigating the 8% Diarization Noise Diarization errors mean the speaker label alone cannot be trusted for commitments (e.g., if Bob says "I'll take that," but the audio engine tags Alice). * The LLM prompt must cross-reference explicit verbal attribution ("Jane, can you take that?" $\rightarrow$ "Sure") against semantic context rather than relying exclusively on transcript speaker IDs. * When semantic confidence falls below threshold ($\tau < 0.75$), flag the action item as Unassigned rather than guessing.

#### Step 4: External Dependencies (Out-of-Room Third Parties) If a meeting participant says: "We need legal to review this, I'll have Sarah from Legal look at it," Sarah was not present and cannot be assigned the task. * Enforcement: The system assigns the task to the speaker (the in-room proxy) as: Assignee: Current Speaker (`[Speaker Name]`) Task: "Coordinate with Sarah (Legal) to review [X]." * If the speaker cannot be deduced with high confidence, set `assignee = UNASSIGNED` with a note: `External dependency: Sarah (Legal)`.

---

3. Product & Feature Specifications

3.1 LLM Processing Pipeline

Within 60 seconds of call termination, the processing service executes the following workflow:

``` [Call Ended] | v [Fetch Clean Transcript + Telemetry EAR] | v [LLM Context Construction] ├── System Prompt & Guardrails ├── Strict EAR List (Names + Emails + IDs) └── Diarized Transcript | v [Structured Output JSON Generation] ├── Executive Summary (Paragraph, <= 150 words) ├── Key Decisions (Bullet points) └── Action Items (Array: Task, Assignee ID, Due Date, Evidence Quote) | v [Deterministic Post-Processing Validator] ├── Cross-reference every Assignee ID against EAR └── Overwrite invalid assignments to "UNASSIGNED" | v [Summary Artifact Created & Persisted] ```

3.2 Action Item Post-Processing Validation Logic

Before writing to the database, run a deterministic verification function:

```python def validateactionitems(extractedactions, eligibleassigneeroster): validatedactions = [] for action in extractedactions: assigneeid = action.get("assigneeid") # Validation: Assignee must explicitly exist in the EAR if assigneeid and assigneeid in eligibleassigneeroster: action["status"] = "assigned" action["assignee"] = eligibleassigneeroster[assigneeid] else: # Fallback for hallucinated or external assignments action["status"] = "unassigned" action["assignee"] = None action["flag"] = "ASSIGNEENOTINROOM" validatedactions.append(action) return validated_actions ```

---

4. Sales Request Resolution: Auto-Email Notifications

The Conflict

Sales requested that action items be emailed automatically to assignees immediately when the meeting ends. However, because speaker labels fail ~8% of the time, immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

The Solution: "Grace Period Auto-Dispatch"

We satisfy Sales' need for immediate execution while protecting data integrity through a time-delayed trigger and host bypass mechanism.

``` [Meeting Ends] │ ├─► AI Generates Artifacts (Target: <60 seconds) │ ├─► Start 5-Minute Grace Period │ │ │ ├─► Host edits/approves? ──► Send instantly with updates │ │ │ └─► Timer expires? ────────► Auto-dispatch without host intervention │ └─► (Workspace Setting Override: "Instant Send" with 0-min delay available) ```

1. Host Review Window (Default): * A 5-minute countdown starts the moment the summary finishes processing. * The meeting host gets a browser/desktop push notification: "Your Huddle summary is ready. Auto-sending in 5:00 minutes. Click to review." * If the host makes edits or clicks "Send Now", the countdown ends and emails dispatch immediately. * If the host takes no action, the emails dispatch automatically at the 5-minute mark. 2. Sales/Workspace Configuration Flag: * Enterprise Admins can toggle the grace window per team (e.g., Sales team can set Grace Period to `0 minutes` for true instant dispatch; default for other workspaces is `5 minutes`). 3. Safety Fallback: Action items marked as `UNASSIGNED` are omitted from individual assignment emails and instead bundled exclusively into the primary meeting recap sent to the host and organizer.

---

5. User Experience & Design Specs

5.1 Post-Meeting In-App Summary Tab

Located inside the Huddle desktop and web apps under the meeting historical detail view.

``` +-----------------------------------------------------------------------+ | Meeting Summary: Product Architecture Sync [Share] [Edit]| | Today at 10:00 AM • 42 mins • 6 Attendees | +-----------------------------------------------------------------------+ | EXECUTIVE SUMMARY | | The team finalized the data migration strategy for Q3. Key risks | | around API rate limiting were addressed by introducing Redis caching.| +-----------------------------------------------------------------------+ | DECISIONS MADE | | • Chose Redis over Memcached for distributed cache. | | • Deferred mobile client updates to Sprint 44. | +-----------------------------------------------------------------------+ | ACTION ITEMS Auto-sending in [ 04:12 ] [x]| | [ + Add Action Item]| | [x] Task: Implement Redis cache layer | | Assignee: [ (Avatar) David Chen v ] Due: [ Friday v ] | | Source: "David: I can set up the cluster by end of week." | | | | [ ] Task: Verify Legal compliance on data retention | | Assignee: [ Unassigned v ] Due: [ Set Date v ] | | ! Flag: External dependency (Sarah / Legal was not present) | +-----------------------------------------------------------------------+ ```

5.2 Key UI Components

1. Assignee Dropdown Menu: * Filtered exclusively to individuals in the $EAR$. * Section 1: Authenticated Attendees (1-click select). * Section 2: Ad-Hoc / Guest Attendees (Displays `[Guest Name] - prompt for email`). * Option to select "Unassigned". * Explicitly disallows searching the global company directory within the "Assignee" field to prevent violating the non-attendee constraint. 2. Context Tooltip ("Source"): * Hovering over an action item shows the transcript snippet that triggered the assignment, allowing the host to quickly verify ambiguous attributions caused by diarization errors.

---

6. Email Dispatch Specs (Sales Notification System)

6.1 Individual Assignee Notification

  • To: Assignee email
  • Reply-To: Host email
  • Subject: `[Action Required] Action item from: {Meeting Title}`

``` Hi {Assignee First Name},

The following action item was assigned to you during {Meeting Title}:

Task: {Task Description} Due Date: {Due Date or "Not specified"} Context: "{Relevant transcript excerpt}"

View full meeting notes and recording: {Link to Huddle Meeting Record}

Did the AI get this wrong? Click here to reassign or remove: {One-Click Edit Link} ```

---

7. Telemetry, Edge Cases & Error Handling

Scenario / Edge CaseFailure Mode / RiskEngineering Resolution
:---:---:---
Multi-person room (Conference hardware)Single client IP/audio line represents 4 people; 1 is calendar-invited, 3 are walk-ins.PRE detects conference room profile. System assigns task to the Conference Room account by default, and sets an inline warning: "Assigned to hardware room. Click to reassign to an individual."
Cross-talk Diarization FlipsAudio assigns Alice's promise ("I'll ship it") to Bob.LLM applies verification: does speaker assignment align with prior conversation context? If confidence score $< 0.75$, fallback to `UNASSIGNED`.
Short/Accidental JoinsAttendee joined for 12 seconds, dropped, and was not present for decisions.PRE requires continuous session duration $\ge 60\text{ seconds}$ to enter the $EAR$.
Uninvited Guest Missing EmailUninvited guest gets an action item; auto-dispatch fails due to missing destination address.Action item is attributed to guest's display name. System flags: "Cannot send email: No address for [Display Name]." Host is prompted with an input box to provide an email.

---

8. Success Metrics

  1. Attribution Precision: $\ge 98\%$ of assigned action items correctly reflect real commitments made by that person (measured via sampling and low manual-edit rates).
  2. Zero-Hallucination Rate: $100\%$ compliance with the hard constraint: 0 automated emails dispatched to individuals outside the $EAR$.
  3. Dispatch Velocity: 95th percentile of summary and action item deliveries completed within $< 90\text{ seconds}$ post-meeting (or exactly at the expiration of the chosen grace period).
  4. Host Intervention Rate: $< 12\%$ of summaries require manual host corrections during the grace period window, indicating high trust in baseline accuracy.

Claims the judge couldn’t find in the brief

  • Manual note-taking results in lost context, untracked deliverables, and administrative drag.
  • Client telemetry logs exist and can establish physical attendance via connection duration greater than 60 seconds.
  • Authenticated users have resolvable user_id, display name, and email; guest joiners have session IDs and display names and may join via guest links.
  • Huddle has desktop and web apps with a meeting historical detail view.
  • There is a global company directory that could be searched for assignees.
  • Conference room accounts or profiles exist and can be assigned by default.
  • Immediately sending unreviewed emails will blast misattributed commitments to clients and internal executives, eroding trust.

Grades and run details

Decision model 50 · LLM judge 1 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly67%
  • passAddresses the actual decision74%
  • failRespects explicit constraints63%
  • failIdentifies material uncertainty41%
  • failAvoids unsupported claims53%
  • passProduces the required deliverable66%
  • passSpecifies AI failure behaviour99%
  • passSuccess settles trade-offs31%
  • partialProposes tests that could fail67%
Run
Run
#1
API response time
35 s
Submitted
1 Oct 2026

Needs more than a quick edit3 of 9 checks passed

The PRD lacks essential trade-off rules and rollout gates a PM would need. It would require rework to define what success rejection looks like and how to make go/no-go decisions.

Output

Product Requirements Document (PRD): AI-Generated Meeting Summaries & Action Items

1. Document Overview & Objective

1.1 Product Name

Huddle Intelligent Summaries (Internal Codename: Project Cue)

1.2 Target Audience

Engineering Leads, Frontend/Backend Engineers, and Product Designers.

1.3 Objective

Build an automated, post-meeting intelligence feature for Huddle that processes meeting transcripts to extract structured decisions, generate comprehensive summaries, and assign action items to attendees. The feature must account for real-world messy transcript data (speaker errors, uninvited participants) and adhere to strict safety boundaries regarding action item assignments and automated distribution.

1.4 Background & Scenario

Huddle is a video-meeting application currently utilized by 3,000 enterprise and mid-market companies. * Existing Infrastructure: Meetings are automatically recorded and transcribed with basic speaker labeling. Attendee lists are reliably synchronized from the originating calendar invite. * Current State & Pain Point: Users currently rely on manual post-meeting note-taking, leading to dropped tasks, misaligned follow-ups, and fragmented alignment. * Value Proposition: Automating meeting wrap-ups directly inside Huddle saves hours of manual administrative overhead per week4, driving higher platform stickiness and daily active use.

---

2. Core User Stories & Personas

2.1 Personas

  • The Engineering Manager (User): Attends 5–6 back-to-back technical syncs daily. Needs clear, unequivocal records of architectural decisions and precise engineering follow-ups without digging through raw, messy transcripts.
  • The Account Executive / Sales Rep (User): Operates on rapid customer cadences. Requires immediate, zero-friction distribution of post-meeting action items to stakeholders without manual typing.
  • The Engineering & Design Leads (Builders): Require explicit edge-case handling, system constraints, deterministic UI layouts, and scalable pipeline specs to build and ship the feature efficiently.

2.2 Key User Stories

  1. As a meeting participant, I want Huddle to automatically generate a structured summary and action item list the moment a meeting concludes, so I don't have to spend time writing manual recap notes.
  2. As an attendee with assigned tasks, I want clear ownership labels on my action items so there is never any ambiguity about who is driving a deliverable.
  3. As a sales representative, I want my action items emailed to me automatically as soon as the meeting ends, so I can immediately execute client follow-ups.
  4. As a meeting host, I want to review, edit, or override incorrect speaker attributions or bad action assignments before summaries are locked or shared externally.

---

3. Detailed Functional Requirements

3.1 Data Inputs & Pre-Processing Pipeline

The AI summarization engine triggers asynchronously via webhook immediately upon meeting termination (`recording.ended`).

  • Input A: Transcript with Speaker Labels: Full text payload containing timestamped utterances, text strings, and speaker IDs. Known limitation: Speaker labels suffer from an 8% error rate, primarily caused by cross-talk, overlapping dialogue, or multi-person huddle rooms.
  • Input B: Calendar Invite Attendee List: Ground-truth list of officially invited participants (Email, Name, Role).1 Known limitation: 15% of meetings include an uninvited participant (e.g., ad-hoc joiners via shared link).
  • Input C: Meeting Metadata: Duration, title, organization ID, and recording URL.

3.2 Summarization & Extraction Logic (LLM Pipeline)

The backend processing pipeline must execute three logical passes: 1. Meeting Overview Generation: Condense the transcript into a 3–5 sentence executive summary alongside 3 key thematic highlights. 2. Decision Extraction: Isolate explicit agreements, architectural choices, or policy resolutions made during the call. Format as an bulleted list of declarative statements (e.g., "Decided to migrate the primary auth service to OAuth2 by Q3"). 3. Action Item Extraction & Assignment Engine: * Extract explicit tasks, owners, and deadlines. * Crucial Constraint (Safety Guardrail): The engine must never assign an action item to someone who was not present in the meeting. * Resolution Strategy for Uninvited Attendees: If an uninvited participant (who joined without a calendar invite) is identified as the owner of a task, the system must cross-reference their presence against the audio/transcript participant logs. If verified present, they are added to a dynamic "Extended Participants" list for assignment validation. If their identity cannot be verified or they were absent, the task defaults to unassigned or falls back to the meeting host with a UI warning flag2.

3.3 Sales Request: Automated Instant Emailing

  • Requirement: Sales teams have requested that action items be emailed automatically to their owners the second a meeting ends.
  • Product Specification & Guardrail: Due to the 8% speaker label error rate and potential hallucination risks, fully autonomous external emailing on meeting close creates high business risk.
  • Implementation Design:
  • Implement an Opt-In Workspace Policy configured by organization admins.
  • For workspaces with "Auto-Send Sales Recaps" enabled, summaries and assigned action items are dispatched via email to internal organization members immediately upon processing completion.3
  • For external participants or non-opt-in workspaces, a 10-minute "Grace Period / Review Window" notification banner appears in the Huddle desktop client, allowing hosts to review and edit assignments before dispatch.

---

4. User Experience & Interface (UX/UI) Specifications

4.1 Post-Meeting View (Huddle Web & Desktop App)

Upon meeting conclusion, a new tab labeled "AI Summary" appears alongside the existing "Transcript" and "Recording" tabs within the meeting details drawer.

``` +-------------------------------------------------------------------+ | Huddle Meeting: Q3 Architecture Sync [Share] [Export]| +-------------------------------------------------------------------+ | Tabs: [Overview] [Transcript] [Recording] | |-------------------------------------------------------------------| | EXECUTIVE SUMMARY | | The team reviewed the Q3 scaling bottlenecks. Consensus was | | reached on moving auth infrastructure to OAuth2. Database latency | | issues will be mitigated via Redis caching layers. | | | | KEY DECISIONS | | • Migrate core auth service to OAuth2 by August 15. [Edit] | | • Adopt Redis cluster for session state caching. [Edit] | | | | ACTION ITEMS | | [x] Investigate OAuth2 token expiry edge cases | | Owner: Sarah Jenkins (Calendar Invite) | Due: Next Tuesday | | [ ] Set up staging Redis cluster | | Owner: Unassigned ⚠️ (Verify Attendee) | Due: Friday | | | | [Send Sales Recap Email Now] [Approve & Lock] | +-------------------------------------------------------------------+ ```

4.2 Interactive States & Edge-Case UI Handling

  • Unassigned / Unverified Owner State: If the model detects an action item owner who was an uninvited participant or has an ambiguous speaker label, render a yellow warning badge: `⚠️ Unverified Owner — Click to Reassign`.
  • Inline Editing: Every decision and action item card features an inline pencil icon allowing hosts or assignees to edit text, change owners via a searchable dropdown of confirmed attendees, or delete hallucinated items.

---

5. Technical Architecture & Engineering Guidelines

5.1 Pipeline Flow & Data Schema

  1. Webhook Listener: Receives `meeting.ended` event containing metadata and storage URIs.
  2. Context Aggregator: Merges calendar invite list (`invitedusers[]`) with actual audio track speaker diarization outputs (`detectedspeakers[]`).
  3. LLM Structured Extraction: Prompt engineering optimized for JSON output schema enforcement (using strict JSON mode / function calling):

```json { "meetingid": "string", "summary": "string", "decisions": ["string"], "actionitems": [ { "task": "string", "assigneeemail": "string | null", "assigneename": "string", "isverifiedattendee": boolean, "due_date": "ISO-8601 | null" } ] } ```

5.2 Error Handling & Fallbacks

  • Speaker Label Inaccuracy (8% Error Rate): Mitigation via UI feedback loops. If an assignee name has low confidence score (< 0.85) from the diarization model, flag the action item with a review indicator in the UI.
  • Uninvited Participant Handling (15% Occurrences): The validation layer must query the roster of active audio streams. If an email is missing from the calendar invite but present in audio metadata, tag them as `guest_participant` and permit assignment only if their voiceprint or explicit self-identification ("Hi, this is Dave from Finance joining late") is logged in the transcript text. Otherwise, fail-safe to unassigned.

---

6. Non-Functional Requirements & Constraints

  • Latency: Summary generation and database write must complete within < 120 seconds of meeting termination for a standard 45-minute meeting.
  • Security & Privacy: Meeting transcripts and summaries must be encrypted at rest (AES-256) and in transit (TLS 1.3). Customer data must never be used to train foundational models across tenant boundaries.
  • Reliability: 99.9% pipeline uptime. In the event of an LLM timeout or parsing failure, the system must gracefully fall back to displaying raw transcripts with an error banner stating "AI summary generation delayed. Click to retry."

---

7. Metrics & Success Criteria (KPIs)

  1. Adoption Rate: % of completed meetings where users view, edit, or export the AI summary (Target: > 40% within 30 days post-launch).
  2. Action Item Completion Rate: Increase in marked-complete action items compared to legacy manual workflows.
  3. Accuracy & Correction Rate: Frequency of manual edits made to action item owners and speaker assignments (Target: < 5% manual override rate, validating low false-positive rates for uninvited attendee checks).
  4. Sales Workflow Efficiency: Reduction in post-call administrative turnaround time reported by enterprise sales teams using the automated email dispatch feature.

What a PM had to fix

  1. 1Constraint missedRestore the constraintStart again

    Source: ConstraintSummaries must never assign an action to someone who was not in the meeting.

    What we’d changeBuild the list of possible owners from who actually joined, not who was invited, and enforce it in code on generated owners, manual edits and sending. A model-set is_verified_attendee flag can't guarantee it.

  2. 2Constraint missedRestore the constraintSubstantial rework

    What we’d changeLeave items with an unverifiable owner unassigned. Handing them to the host invents responsibility nobody agreed to.

  3. 3Constraint missedRestore the constraintSubstantial rework

    Source: Known limitsSpeaker labels are wrong about 8% of the time

    What we’d changeDon't auto-send. With speaker labels wrong 8% of the time, action items would reach the wrong owners before anyone checks them. Keep organiser review before any email.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent the time saved as a goal to measure, not a known benefit.

Claims the judge couldn’t find in the brief

  • Huddle is used by 3,000 enterprise and mid-market companies.
  • Huddle has a Huddle Web & Desktop App.

Grades and run details

Decision model 50 · LLM judge 4 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly52%
  • passAddresses the actual decision42%
  • failRespects explicit constraints43%
  • partialIdentifies material uncertainty24%
  • failAvoids unsupported claims40%
  • passProduces the required deliverable53%
  • passSpecifies AI failure behaviour93%
  • partialSuccess settles trade-offs57%
  • partialProposes tests that could fail58%
Artefacts
Run
Run
#1
Time to output
15 s
Submitted
24 Sept 2026

Usable with a quick edit5 of 9 checks passed

Yes. A PM could quickly add measurement windows to the metrics and rephrase the unsupported problem-statement claims without redoing any structural or reasoning work.

Output

PRD: AI Meeting Summaries and Action Items

Product: Huddle Audience: Engineering and Design leads Status: Proposed MVP Launch decision: Ship reviewed summaries first; defer immediate, unreviewed owner emails.

1. Problem and opportunity

Huddle serves 3,000 companies. Meetings are recorded and transcribed with speaker labels, but attendees currently write their own notes afterward. This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

We will generate a concise meeting summary containing decisions and action items, with evidence linking each item to the recording.

Two known limitations shape the design:

  • Speaker labels are wrong approximately 8% of the time, especially during overlapping speech and shared-room participation.
  • Approximately 15% of meetings include someone absent from the calendar invite.

Therefore, neither speaker labels nor calendar invitations are sufficient evidence for assigning work.

Non-negotiable requirement: An action must never be assigned to someone who did not attend the meeting. When attendance or identity cannot be established, leave the action unassigned.

2. Goals and non-goals

Goals

  • Reduce time spent producing post-meeting notes.
  • Capture decisions accurately, distinguishing decisions from proposals.
  • Extract actionable commitments with owners and deadlines when supported.
  • Include verified, uninvited attendees without treating all invitees as attendees.
  • Make every decision and action inspectable and correctable.
  • Deliver owner notifications only after safe assignment and publication.

Non-goals for MVP

  • Live summaries or action assignment during meetings.
  • Assigning work to absent colleagues, even if requested during the meeting.
  • Inferring deadlines or assigning owners based on job title.
  • External task-system integrations.
  • Automatically publishing unreviewed summaries.
  • Resolving identity using voice biometrics.

3. Users and primary workflow

Reviewer: Meeting organizer, or an explicitly designated attendee. Reviews and publishes the summary.

Attendee: Reads the summary, follows evidence links, and reports corrections.

Action owner: Receives published actions assigned to them.

Workflow:

  1. Meeting ends; Huddle completes transcript processing.
  2. Huddle reconciles actual attendance and generates a private draft.
  3. Reviewer checks decisions, unresolved ownership, and attendance exceptions.
  4. Reviewer edits and publishes.
  5. Huddle makes the summary available to authorized viewers and, where enabled, emails owners their approved actions.

If no eligible reviewer is available, the draft remains unpublished. A non-attending organizer may designate an attending reviewer but cannot resolve attendance exceptions themselves.

4. Attendance and identity requirements

Engineering must establish an attendance ledger separate from the calendar invite.

Each record contains:

  • Meeting ID and stable person ID, where resolved.
  • Join/leave intervals and participant-session IDs.
  • Identity source: authenticated join, verified guest identity, or shared-room roster.
  • Verification state and evidence.
  • Reviewer confirmation details, where required.

Calendar invitees are suggestions for identity matching, not proof of attendance. A directory search result is also not attendance evidence.

Assignment eligibility

An owner must map to a unique person with verified attendance in that meeting:

  • Authenticated individual join: Eligible when attendance telemetry records the person’s session.
  • Guest join: Eligible only after identity is resolved and verified; a display name alone is insufficient.
  • Shared meeting room: The room endpoint proves room participation, not individual presence. A participating reviewer must explicitly confirm each person physically present before they become eligible.
  • Uninvited attendee: Eligible through the same verification rules; invitation status is irrelevant.
  • Invited but absent person: Ineligible.

If telemetry is unavailable, identity is ambiguous, or attendance is disputed, block assignment. Keep the action unassigned with an explanation.

The owner picker must show only eligible attendees. Every write path—including generation, manual edits, APIs, and publication—must enforce eligibility server-side. Model confidence must never override this rule.

These controls cannot independently observe who is physically in a room. If trustworthy roster confirmation is unavailable, shared-room individuals remain ineligible rather than weakening the constraint.

5. Summary and extraction requirements

The draft has four sections:

  1. Overview: Two to four sentences describing the meeting’s purpose and outcome.
  2. Decisions: Clear statements of agreed outcomes.
  3. Action items: Task, owner or “Unassigned,” deadline if explicit, and status.
  4. Needs review: Uncertain decisions, ownership, identity, or attendance.

Each decision and action must include transcript-span references and recording timestamps. Preserve the original supporting text even when the reviewer edits the final wording.

Decisions

Capture explicit agreement or a clearly accepted conclusion. Do not convert brainstorming, recommendations, or unresolved debate into decisions. If discussion reverses an earlier decision, represent the final outcome and retain supporting evidence.

Actions

Extract a concrete deliverable or next step. Record a deadline only when explicitly stated; normalize relative dates using the meeting’s date and timezone.

Examples:

  • “I’ll send the revised deck tomorrow.” → Task with a candidate owner, subject to attribution review.
  • “Someone should update the deck.” → Unassigned action.
  • “Ask Priya to update it,” when Priya was absent → Unassigned action; note that an absent person was mentioned.
  • “We could revisit this next quarter.” → Not an action without a clear commitment.

Speaker labels provide evidence, not authority. Overlap, shared-room speech, ambiguous pronouns, and conflicting commitments must trigger ownership review. Candidate owners are draft suggestions, not assignments, until approved.

6. Review experience

The summary page opens in Draft state with a clear “Review before sharing” banner.

Design must provide:

  • Editable decision and action cards.
  • Evidence links opening the relevant transcript and recording segment.
  • Visible distinction between verified attendance and uncertain attribution.
  • “Unassigned” as a first-class owner option.
  • A dedicated queue for identity and attendance exceptions.
  • A publish preview showing recipients and notification settings.

For each assigned action, the reviewer must approve the owner. Bulk approval is permitted for straightforward items, but flagged attribution requires item-level review.

Attendance verification and ownership approval are separate controls: confirming that someone attended does not prove they accepted a task.

Publication is blocked if any assigned owner is ineligible. Unassigned actions and unresolved decision candidates do not block publication; uncertain decisions must be omitted from the final Decisions section or explicitly labeled unresolved.

7. Email policy and Sales request

Sales requests automatic emails immediately after each meeting. MVP will not send unreviewed action assignments. Given the known speaker errors, immediate delivery would spread incorrect ownership before correction.

Instead:

  • Notify the reviewer when the draft is ready.
  • Default owner emails to off; enable through workspace settings.
  • When enabled, automatically email approved owners after publication.
  • Send each owner only their actions, meeting title, deadlines, and an authorized summary link.
  • Never email a calendar invitee solely because they were invited.
  • Do not send to guests without a verified address and permitted access.

Email dispatch must recheck eligibility and publication version. Use idempotency keys to prevent duplicate sends. Published ownership changes generate a correction for the previous recipient and an updated notification for the new eligible owner.

Future immediate automation requires a separate launch decision and evidence that it preserves the attendance constraint.

8. Technical implementation

Use an asynchronous pipeline:

`Meeting ended → Transcript ready → Attendance reconciled → Draft generated → Reviewed → Published → Notifications`

Proposed records:

  • `Summary`: meeting ID, transcript version, state, generated content, revision.
  • `Decision`: text, evidence spans, review state.
  • `Action`: task, proposed owner, approved owner, deadline, evidence, review state.
  • `AttendanceRecord`: identity, evidence, verification state.
  • `AuditEvent`: actor, change, timestamp, revision.

Generation returns schema-validated structured data. Treat transcript content as untrusted input: instructions spoken during a meeting must not alter system rules, access controls, or recipient selection.

Owner writes and publication run through a centralized eligibility validator. Notification jobs check the current revision before sending. Attendance corrections invalidate affected assignments and cancel queued notifications. Maintain an audit trail rather than silently rewriting history.

9. Reliability, privacy, and access

Target draft availability: within five minutes of transcript completion for 95% of supported meetings.

Generation failures must leave recordings and transcripts usable, show retry status, and send no owner emails.

Drafts are visible only to reviewers. Published summaries inherit meeting-content access rules; attendance alone does not grant access. Recording evidence links must enforce existing permissions. Apply workspace retention and deletion policies to summaries, attendance evidence, and generated artifacts.

10. Success metrics and acceptance

Proposed launch gates

  • Zero assignments to ineligible attendees across automated invariant tests.
  • Zero emails to ineligible owners in the pilot; any confirmed incident pauses rollout.
  • At least 90% of published decisions require no substantive reviewer correction.
  • Median review time below three minutes for meetings under one hour.
  • At least 30% less self-reported note-taking time in the pilot.

Measure owner-correction rate separately from attendance violations: a wrong attending owner is still a quality failure.

Required acceptance cases

Test absent invitees, verified uninvited guests, shared rooms, overlapping speakers, duplicate names, failed identity resolution, missing telemetry, absent people named in requests, transcript reprocessing, attendance revoked before dispatch, and repeated notification retries.

In every uncertain-attendance case, expected behavior is an unassigned action and no owner email.

11. Rollout and dependencies

Start with internal meetings, then an opt-in pilot across varied company sizes and shared-room usage. Expand only after launch gates pass.

Before implementation, confirm availability of participant-session telemetry, guest identity verification, room-roster confirmation, reviewer permissions, and verified email addresses. If these dependencies are missing, ship summaries with unassigned actions for affected participants rather than substituting calendar membership as attendance proof.

Claims the judge couldn’t find in the brief

  • This creates duplicated effort, inconsistent records of decisions, and missed follow-through.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly12%
  • passAddresses the actual decision98%
  • passRespects explicit constraints68%
  • passIdentifies material uncertainty87%
  • failAvoids unsupported claims32%
  • passProduces the required deliverable83%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs35%
  • partialProposes tests that could fail73%
Run
Run
#1
API response time
47 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Addresses the actual decisionMixedRightRight
Gemini 3.8 Flash · API

It chooses a grace-period auto-dispatch answer but does not clearly state what result or condition would change that decision.

Gemini 3.5 Flash-Lite · Gemini

The PRD commits to a clear design, addressing the Sales request with an opt-in auto-email and a review window, and is unambiguous for engineering and design leads.

GPT-6.1 Sol · API

The output commits early to 'Ship reviewed summaries first; defer immediate, unreviewed owner emails' and states that future immediate automation requires evidence that it preserves the attendance constraint.

Respects explicit constraintsWrongMixedRight
Gemini 3.8 Flash · API

It violates the prohibition on auto-emailing without review by dispatching after five minutes even if no one reviews, and it likely exceeds the 1,000–1,500 word limit.

Gemini 3.5 Flash-Lite · Gemini

The constraint that summaries never assign actions to non-attendees is enforced via verification and fallback to unassigned/host; the PRD also fits the requested form and length.

GPT-6.1 Sol · API

The output enforces the non-assignment constraint through an attendance ledger, eligibility checks, server-side validation, and reviewer approval; it also addresses the Sales request by deferring auto-email until after review.

Identifies material uncertaintyWrongWrongRight
Gemini 3.8 Flash · API

It names known error rates but does not bound the material unknowns that could change the auto-email decision or specify what would change the call.

Gemini 3.5 Flash-Lite · Gemini

It mentions error rates but does not specify what conditions would change the design decision (e.g., an override-rate threshold that would trigger disabling auto-email).

GPT-6.1 Sol · API

The output names specific unknowns such as shared-room physical attendance, missing telemetry, and ambiguous identity, and says how they would be resolved (reviewer confirmation, unassigned actions) and what would change the call (evidence for future automation).

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

Although it is a detailed PRD, it is too long for the requested range and contains major unsupported assumptions that engineering and design could not act on without re-checking.

Gemini 3.5 Flash-Lite · Gemini

The output is a structured PRD with functional specs, UI mockups, and technical architecture, within 1,000–1,500 words, and is usable by engineering and design leads.

GPT-6.1 Sol · API

The output is a complete PRD aimed at engineering and design leads, covering decisions, actions, attribution, editing, and failure behaviour, and it appears to be within the requested 1,000-1,500 word range.

Success settles trade-offsMixedWrongMixed
Gemini 3.8 Flash · API

It gives metric targets but no measurement time frame and no explicit rule for which goal wins when metrics conflict.

Gemini 3.5 Flash-Lite · Gemini

No explicit trade-off rule is given (e.g., accepting lower coverage to keep assignment precision above a stated level), and not all success metrics have targets and time frames.

GPT-6.1 Sol · API

The success metrics have numeric targets but no explicit time frame (e.g., pilot duration), and the output lacks a single clear trade-off rule stating which goal wins when precision and coverage conflict, though unassigned-over-misassignment is implied.

All got wrong 3

Uses the supplied evidence correctlyWrongWrongWrong
Gemini 3.8 Flash · API

It invents current systems and data such as telemetry logs, guest session IDs, desktop/web historical views, global directory, and conference room accounts that are not in the supplied context.

Gemini 3.5 Flash-Lite · Gemini

The output claims Huddle is used by 'enterprise and mid-market companies' and has a 'Huddle Web & Desktop App', neither of which is in the supplied context.

GPT-6.1 Sol · API

The output presents 'This creates duplicated effort, inconsistent records of decisions, and missed follow-through' as a current-situation fact, but it is not in the supplied context and does not follow by arithmetic.

Avoids unsupported claimsWrongWrongWrong
Gemini 3.8 Flash · API

It presents several assumptions and forecasts as established facts, including telemetry availability, existing UI surfaces, and inevitable trust damage from immediate emails.

Gemini 3.5 Flash-Lite · Gemini

It presents 'enterprise and mid-market' and the existence of a desktop client as established facts without backing from the supplied evidence.

GPT-6.1 Sol · API

The output states 'This creates duplicated effort...' and 'immediate delivery would spread incorrect ownership' as established facts without labelling them as inferences or forecasts, although the supplied context only suggests these outcomes.

Proposes tests that could failWrongWrongWrong
Gemini 3.8 Flash · API

Its gates and metrics lack measurement windows and clear actions for each pass/fail outcome.

Gemini 3.5 Flash-Lite · Gemini

No tests, gates, or kill criteria with numeric thresholds, measurement windows, and consequents are proposed.

GPT-6.1 Sol · API

The launch gates and acceptance cases have numeric thresholds and some actions, but they lack measurement windows long enough to read out the effects, and not every gate specifies the action if the threshold is not met.

All got right 1

Specifies AI failure behaviourRightRightRight
Gemini 3.8 Flash · API

It specifies low-confidence unassignment, deterministic validation, flags, editing, and missing-email handling.

Gemini 3.5 Flash-Lite · Gemini

It defines UI warning states for unverified owners, low confidence, and fallback to raw transcript on LLM timeout.

GPT-6.1 Sol · API

The PRD explicitly specifies what happens when identity or attendance is uncertain (unassigned actions, no emails), when reviewers are unavailable (draft remains unpublished), and when generation fails (no owner emails, recordings usable).

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD