Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for AI-generated meeting summaries that list the decisions made and assign action items to attendees. The PRD is for the engineering and design leads who will build it: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given4 items: Scenario, Known limits, Constraint, Request from Sales
ScenarioHuddle is a video-meeting app used by 3,000 companies. Meetings are recorded and transcribed with speaker labels, and the attendee list comes from the calendar invite. Today people write their own notes after the call.
Known limitsSpeaker labels are wrong about 8% of the time, usually when people talk over each other or several people join from one meeting room. About 15% of meetings include someone who joined without a calendar invite.
ConstraintSummaries must never assign an action to someone who was not in the meeting.
Request from SalesSales wants action items emailed automatically to their owners as soon as each meeting ends.
What a strong answer doesThe answer key the graders mark against

Specifies attribution confidence, editing, and what happens on mis-assignment; keeps a person reviewing before action items go out.

Critical failures (cap the score at 40)
  • Assigns actions to non-attendees
Case

Mid-level · v1.10 · synthetic · AI product, productivity

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 9 checks passed

The PRD is mostly usable but assumes session join/leave logs exist and repeats unsourced claims about note quality and room devices, so the evidence would need re-checking. Some success criteria also lack exact thresholds and actions.

Output

PRD: AI Meeting Summaries with Decisions and Action Items

Status: Draft for engineering and design leads Product: Huddle

1. Problem

Huddle records and transcribes meetings with speaker labels, but attendees still write their own notes afterward. Notes are late, inconsistent, or never written, so decisions get lost and action items have no clear owner. We will generate a summary after each meeting that lists decisions and proposes action items with owners.

2. Goals and non-goals

Goals - Produce a summary within 5 minutes of meeting end: decisions, action items (task, owner, due date if stated), and open questions. - Never assign an action item to someone who was not in the meeting (hard constraint, see §4). - Keep the host's review to under 60 seconds for a typical meeting.

Non-goals (v1) - Fully automatic emailing of action items to owners (see §3). - Cross-meeting task tracking or integrations with Jira, Asana, and similar tools. - Summaries for meetings without a transcript.

3. Decision on the Sales request

Sales asked for action items to be emailed automatically to owners the moment a meeting ends. We are not building that for v1. Instead, we will send the draft summary to the host at meeting end. The host can send it to owners with one click, and owners are emailed on send.

Reasons: - Speaker labels are wrong about 8% of the time. An automatic email would sometimes tell a real person they own a task they never agreed to. Because it goes out with no human check, the error becomes visible to a customer's colleagues. - Email is irreversible, so a wrong or non-attendee assignment cannot be quietly fixed afterward. - The underlying Sales need is that owners hear about their tasks quickly and nothing falls through the cracks. A host prompt at meeting end with one-click send meets that need at the cost of a few seconds.

Path to automation: Auto-send will be an opt-in, per-workspace setting, considered after launch once the gates in §9 are met. It would send only high-confidence, presence-verified items and leave everything else for host review. Engineering should build the send pipeline so that auto-send is later a policy change rather than a rebuild.

4. The hard constraint: who can own an action item

Rule: An action item's owner must be a participant who was actually in the meeting when the item was discussed. Anyone else is ineligible.

This rule cannot rely on the calendar invite, and it cannot rely on the model behaving well.

Attendance source. The eligible-owner roster is built from Huddle's session join/leave logs, not from the calendar invite. This handles both failure modes: - Invitees who never joined are not on the roster. - The roughly 15% of meetings with uninvited joiners include those people, matched to a user account when possible.

Presence at time of item. A person is eligible for an item only if they were in the session at the timestamp of the relevant utterance. Someone who left at minute 10 cannot own an item from minute 40.

Enforcement is structural, not prompt-based. 1. The extraction step receives the roster as a list of participant IDs and must return an owner ID from that list or `null`. 2. A deterministic validator runs after the model. It rejects any owner ID not on the roster or not present at the item's timestamp, and it converts that item to unassigned. 3. The UI and email layers accept only validated owner IDs. There is no free-text owner field on the model output path. 4. Names mentioned in the transcript who were not present ("have Dana look at it") produce an unassigned item. The text may keep the speaker's wording, but Dana gets no assignment and no email.

Shared rooms. A room device usually appears as one participant, so individuals in the room are not in the presence log. We cannot verify that a named person is in the room. Such items become unassigned by default, and the host can attest who was in the room during review (§6). That attestation counts as human-verified presence.

Uninvited joiners without an account (guests) can be eligible owners, shown by display name. They receive email only if the host supplies an address during review. We never guess addresses.

5. Handling speaker-label errors

The 8% label error rate is concentrated in crosstalk and shared rooms. Eligible-owner enforcement does not solve it, because a mislabel can still pick the wrong person who was present. Mitigations:

SignalHandling
Owner named explicitly ("Priya, can you send the deck?") and Priya is present, followed by assent or no objectionHigh confidence. Name-based assignment does not depend on speaker labels.
Self-assignment ("I'll send it")Depends on the label. High only if the segment is not in a crosstalk or shared-room span and label confidence is above threshold. Otherwise Medium.
Assignment inferred from context, or the segment overlaps other speakersLow. Default to unassigned with a suggested owner the host can accept.
Speaker is a shared-room deviceNever treated as a person. Unassigned unless a named individual is explicitly addressed and verified.

Confidence tiers (High/Medium/Low) are stored on each item, shown to the host in review, and used later to gate auto-send. Thresholds are set empirically during the evaluation in §9. The tiers themselves are a v1 requirement.

6. User experience

Host flow 1. The meeting ends and the host gets a notification: "Summary ready to review." 2. The review screen lists: - Decisions as short statements, each linked to the transcript moment. - Action items with the task, owner chip, due date if stated, and a confidence indicator. - Unassigned items grouped at the top of the actions list, so gaps are visible. 3. The host can edit text, reassign from a picker limited to eligible people, delete items, or add room attendees ("Who was in the room?"). Adding room attendees makes them eligible for items in the room segments. 4. Send emails each owner their own items plus the decisions. Unassigned items go to no one.

Design requirements - Every item has a "jump to source" link that plays the transcript segment. Reviewing should be verification, not re-listening. - The owner picker never offers non-attendees. If the host wants to assign to an absent person, show a separate option, "Send as a follow-up (not from the meeting)," which is a different item type with clearly different wording and no "you agreed to" language. Design to decide whether this ships in v1; the default is that it does not. - Low-confidence owners appear as suggestions ("Suggested: Sam") and never as assignments. - Email copy says "Action item from Meeting name, date", includes the source excerpt, and offers a "This isn't mine" link that notifies the host. This gives owners a correction path. - If the host does not review, send a reminder at 24 hours. Nothing is sent to owners without host action in v1.

7. Functional requirements

  1. Roster service: builds the per-meeting roster with join/leave intervals from session logs and links to user accounts where possible.
  2. Extraction: from the transcript and roster, produce decisions, action items, and unassigned items. Each action item has task text, owner ID or null, due date or null, source timestamps, and a confidence tier.
  3. Validator: enforces §4 as code with unit tests, independent of the model, and logs every rejection.
  4. Review UI: as described in §6, on web and desktop. Mobile is view-only for v1.
  5. Delivery: per-owner emails on host send, with a record of who was emailed. Sends are idempotent, and edits after sending trigger an explicit "send update" action.
  6. Audit log: stores for each item the model output, validator result, host edits, and send events.
  7. Admin setting: workspace-level toggle for summaries, plus retention aligned with existing transcript retention.

8. Edge cases to specify and test

  • Invitee never joined but is mentioned as owner.
  • Uninvited joiner takes an action; the guest has no account.
  • Person leaves and rejoins (multiple intervals).
  • Two people share a first name; ownership must resolve by roster, and ambiguity means unassigned.
  • Whole meeting in a shared room (all items default unassigned until the host attests).
  • Multiple languages, or a transcript with large gaps.
  • Nobody commits to anything: show "No action items detected" rather than inventing some.

9. Success metrics and launch gates

Launch gates (must all pass) - Zero non-attendee assignments across a labeled evaluation set of at least 500 meetings, deliberately oversampled for uninvited joiners, shared rooms, no-show invitees, and crosstalk. Any production occurrence afterward is a Sev-1 incident. - Among High-confidence items, wrong-attendee assignment rate of 3% or less (target, to be validated against the evaluation set; propose a stricter figure if data supports it). - Median summary delivery within 5 minutes of meeting end.

Post-launch metrics - Share of meetings where the host sends the summary within 1 hour. - Host edit rate per item, and owner "This isn't mine" rate. - Share of action items left unassigned. Some is healthy, but a high rate means the model is too timid.

Auto-send criteria (future): two consecutive months of High-tier precision above the agreed bar and no non-attendee incidents. Sales and Support are consulted before enabling.

10. Risks and open questions

  • Sales disagreement. Sales may see the host step as friction. We should share §3 with them early and offer a pilot of auto-send with a few design-partner accounts after the gates are met.
  • Hosts skipping review. If sending rates are low, the fix is better reminders and faster review, not removing the check.
  • Presence data quality. Do our session logs reliably capture join/leave for all clients and dial-in users? Engineering to confirm in week 1.
  • Guest privacy. Can we email or display guest names without consent, given customer policies? Legal to advise.
  • Room attestation. Is host attestation adequate for the constraint, or should room participants confirm individually? Recommend host attestation for v1.
  • Follow-up items for absent people. Ship in v1 or defer? Design and Product to decide.

11. Milestones

  1. Weeks 1–2: roster service, presence data audit, evaluation set labeling begins.
  2. Weeks 3–5: extraction, validator, confidence tiers, offline evaluation.
  3. Weeks 4–7: review UI and email delivery, in parallel.
  4. Weeks 8–9: internal dogfood, gate review.
  5. Week 10: limited beta with design-partner customers, including Sales' accounts.

Claims the judge couldn’t find in the brief

  • Huddle has session join/leave logs that can build a per-meeting roster.
  • Notes are late, inconsistent, or never written.
  • A room device usually appears as one participant, so individuals in the room are not in the presence log.
  • Email is irreversible.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly47%
  • passAddresses the actual decision99%
  • passRespects explicit constraints71%
  • passIdentifies material uncertainty97%
  • failAvoids unsupported claims26%
  • passProduces the required deliverable80%
  • passSpecifies AI failure behaviour100%
  • passSuccess settles trade-offs74%
  • partialProposes tests that could fail62%
Run
Run
#1
API response time
49 s
Submitted
29 Sept 2026

Needs more than a quick edit3 of 9 checks passed

The PRD lacks essential trade-off rules and rollout gates a PM would need. It would require rework to define what success rejection looks like and how to make go/no-go decisions.

Output

Product Requirements Document (PRD): AI-Generated Meeting Summaries & Action Items

1. Document Overview & Objective

1.1 Product Name

Huddle Intelligent Summaries (Internal Codename: Project Cue)

1.2 Target Audience

Engineering Leads, Frontend/Backend Engineers, and Product Designers.

1.3 Objective

Build an automated, post-meeting intelligence feature for Huddle that processes meeting transcripts to extract structured decisions, generate comprehensive summaries, and assign action items to attendees. The feature must account for real-world messy transcript data (speaker errors, uninvited participants) and adhere to strict safety boundaries regarding action item assignments and automated distribution.

1.4 Background & Scenario

Huddle is a video-meeting application currently utilized by 3,000 enterprise and mid-market companies. * Existing Infrastructure: Meetings are automatically recorded and transcribed with basic speaker labeling. Attendee lists are reliably synchronized from the originating calendar invite. * Current State & Pain Point: Users currently rely on manual post-meeting note-taking, leading to dropped tasks, misaligned follow-ups, and fragmented alignment. * Value Proposition: Automating meeting wrap-ups directly inside Huddle saves hours of manual administrative overhead per week4, driving higher platform stickiness and daily active use.

---

2. Core User Stories & Personas

2.1 Personas

  • The Engineering Manager (User): Attends 5–6 back-to-back technical syncs daily. Needs clear, unequivocal records of architectural decisions and precise engineering follow-ups without digging through raw, messy transcripts.
  • The Account Executive / Sales Rep (User): Operates on rapid customer cadences. Requires immediate, zero-friction distribution of post-meeting action items to stakeholders without manual typing.
  • The Engineering & Design Leads (Builders): Require explicit edge-case handling, system constraints, deterministic UI layouts, and scalable pipeline specs to build and ship the feature efficiently.

2.2 Key User Stories

  1. As a meeting participant, I want Huddle to automatically generate a structured summary and action item list the moment a meeting concludes, so I don't have to spend time writing manual recap notes.
  2. As an attendee with assigned tasks, I want clear ownership labels on my action items so there is never any ambiguity about who is driving a deliverable.
  3. As a sales representative, I want my action items emailed to me automatically as soon as the meeting ends, so I can immediately execute client follow-ups.
  4. As a meeting host, I want to review, edit, or override incorrect speaker attributions or bad action assignments before summaries are locked or shared externally.

---

3. Detailed Functional Requirements

3.1 Data Inputs & Pre-Processing Pipeline

The AI summarization engine triggers asynchronously via webhook immediately upon meeting termination (`recording.ended`).

  • Input A: Transcript with Speaker Labels: Full text payload containing timestamped utterances, text strings, and speaker IDs. Known limitation: Speaker labels suffer from an 8% error rate, primarily caused by cross-talk, overlapping dialogue, or multi-person huddle rooms.
  • Input B: Calendar Invite Attendee List: Ground-truth list of officially invited participants (Email, Name, Role).1 Known limitation: 15% of meetings include an uninvited participant (e.g., ad-hoc joiners via shared link).
  • Input C: Meeting Metadata: Duration, title, organization ID, and recording URL.

3.2 Summarization & Extraction Logic (LLM Pipeline)

The backend processing pipeline must execute three logical passes: 1. Meeting Overview Generation: Condense the transcript into a 3–5 sentence executive summary alongside 3 key thematic highlights. 2. Decision Extraction: Isolate explicit agreements, architectural choices, or policy resolutions made during the call. Format as an bulleted list of declarative statements (e.g., "Decided to migrate the primary auth service to OAuth2 by Q3"). 3. Action Item Extraction & Assignment Engine: * Extract explicit tasks, owners, and deadlines. * Crucial Constraint (Safety Guardrail): The engine must never assign an action item to someone who was not present in the meeting. * Resolution Strategy for Uninvited Attendees: If an uninvited participant (who joined without a calendar invite) is identified as the owner of a task, the system must cross-reference their presence against the audio/transcript participant logs. If verified present, they are added to a dynamic "Extended Participants" list for assignment validation. If their identity cannot be verified or they were absent, the task defaults to unassigned or falls back to the meeting host with a UI warning flag2.

3.3 Sales Request: Automated Instant Emailing

  • Requirement: Sales teams have requested that action items be emailed automatically to their owners the second a meeting ends.
  • Product Specification & Guardrail: Due to the 8% speaker label error rate and potential hallucination risks, fully autonomous external emailing on meeting close creates high business risk.
  • Implementation Design:
  • Implement an Opt-In Workspace Policy configured by organization admins.
  • For workspaces with "Auto-Send Sales Recaps" enabled, summaries and assigned action items are dispatched via email to internal organization members immediately upon processing completion.3
  • For external participants or non-opt-in workspaces, a 10-minute "Grace Period / Review Window" notification banner appears in the Huddle desktop client, allowing hosts to review and edit assignments before dispatch.

---

4. User Experience & Interface (UX/UI) Specifications

4.1 Post-Meeting View (Huddle Web & Desktop App)

Upon meeting conclusion, a new tab labeled "AI Summary" appears alongside the existing "Transcript" and "Recording" tabs within the meeting details drawer.

``` +-------------------------------------------------------------------+ | Huddle Meeting: Q3 Architecture Sync [Share] [Export]| +-------------------------------------------------------------------+ | Tabs: [Overview] [Transcript] [Recording] | |-------------------------------------------------------------------| | EXECUTIVE SUMMARY | | The team reviewed the Q3 scaling bottlenecks. Consensus was | | reached on moving auth infrastructure to OAuth2. Database latency | | issues will be mitigated via Redis caching layers. | | | | KEY DECISIONS | | • Migrate core auth service to OAuth2 by August 15. [Edit] | | • Adopt Redis cluster for session state caching. [Edit] | | | | ACTION ITEMS | | [x] Investigate OAuth2 token expiry edge cases | | Owner: Sarah Jenkins (Calendar Invite) | Due: Next Tuesday | | [ ] Set up staging Redis cluster | | Owner: Unassigned ⚠️ (Verify Attendee) | Due: Friday | | | | [Send Sales Recap Email Now] [Approve & Lock] | +-------------------------------------------------------------------+ ```

4.2 Interactive States & Edge-Case UI Handling

  • Unassigned / Unverified Owner State: If the model detects an action item owner who was an uninvited participant or has an ambiguous speaker label, render a yellow warning badge: `⚠️ Unverified Owner — Click to Reassign`.
  • Inline Editing: Every decision and action item card features an inline pencil icon allowing hosts or assignees to edit text, change owners via a searchable dropdown of confirmed attendees, or delete hallucinated items.

---

5. Technical Architecture & Engineering Guidelines

5.1 Pipeline Flow & Data Schema

  1. Webhook Listener: Receives `meeting.ended` event containing metadata and storage URIs.
  2. Context Aggregator: Merges calendar invite list (`invitedusers[]`) with actual audio track speaker diarization outputs (`detectedspeakers[]`).
  3. LLM Structured Extraction: Prompt engineering optimized for JSON output schema enforcement (using strict JSON mode / function calling):

```json { "meetingid": "string", "summary": "string", "decisions": ["string"], "actionitems": [ { "task": "string", "assigneeemail": "string | null", "assigneename": "string", "isverifiedattendee": boolean, "due_date": "ISO-8601 | null" } ] } ```

5.2 Error Handling & Fallbacks

  • Speaker Label Inaccuracy (8% Error Rate): Mitigation via UI feedback loops. If an assignee name has low confidence score (< 0.85) from the diarization model, flag the action item with a review indicator in the UI.
  • Uninvited Participant Handling (15% Occurrences): The validation layer must query the roster of active audio streams. If an email is missing from the calendar invite but present in audio metadata, tag them as `guest_participant` and permit assignment only if their voiceprint or explicit self-identification ("Hi, this is Dave from Finance joining late") is logged in the transcript text. Otherwise, fail-safe to unassigned.

---

6. Non-Functional Requirements & Constraints

  • Latency: Summary generation and database write must complete within < 120 seconds of meeting termination for a standard 45-minute meeting.
  • Security & Privacy: Meeting transcripts and summaries must be encrypted at rest (AES-256) and in transit (TLS 1.3). Customer data must never be used to train foundational models across tenant boundaries.
  • Reliability: 99.9% pipeline uptime. In the event of an LLM timeout or parsing failure, the system must gracefully fall back to displaying raw transcripts with an error banner stating "AI summary generation delayed. Click to retry."

---

7. Metrics & Success Criteria (KPIs)

  1. Adoption Rate: % of completed meetings where users view, edit, or export the AI summary (Target: > 40% within 30 days post-launch).
  2. Action Item Completion Rate: Increase in marked-complete action items compared to legacy manual workflows.
  3. Accuracy & Correction Rate: Frequency of manual edits made to action item owners and speaker assignments (Target: < 5% manual override rate, validating low false-positive rates for uninvited attendee checks).
  4. Sales Workflow Efficiency: Reduction in post-call administrative turnaround time reported by enterprise sales teams using the automated email dispatch feature.

What a PM had to fix

  1. 1Constraint missedRestore the constraintStart again

    Source: ConstraintSummaries must never assign an action to someone who was not in the meeting.

    What we’d changeBuild the list of possible owners from who actually joined, not who was invited, and enforce it in code on generated owners, manual edits and sending. A model-set is_verified_attendee flag can't guarantee it.

  2. 2Constraint missedRestore the constraintSubstantial rework

    What we’d changeLeave items with an unverifiable owner unassigned. Handing them to the host invents responsibility nobody agreed to.

  3. 3Constraint missedRestore the constraintSubstantial rework

    Source: Known limitsSpeaker labels are wrong about 8% of the time

    What we’d changeDon't auto-send. With speaker labels wrong 8% of the time, action items would reach the wrong owners before anyone checks them. Keep organiser review before any email.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent the time saved as a goal to measure, not a known benefit.

Claims the judge couldn’t find in the brief

  • Huddle is used by 3,000 enterprise and mid-market companies.
  • Huddle has a Huddle Web & Desktop App.

Grades and run details

Decision model 50 · LLM judge 4 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly52%
  • passAddresses the actual decision42%
  • failRespects explicit constraints43%
  • partialIdentifies material uncertainty24%
  • failAvoids unsupported claims40%
  • passProduces the required deliverable53%
  • passSpecifies AI failure behaviour93%
  • partialSuccess settles trade-offs57%
  • partialProposes tests that could fail58%
Artefacts
Run
Run
#1
Time to output
15 s
Submitted
24 Sept 2026

Needs more than a quick edit2 of 9 checks passed

It needs rework to require human review before auto-email and to replace unsupported attendance data assumptions. It cannot be used with only a quick edit.

Output

PRD: AI-Generated Meeting Summaries and Action Items

Audience: Engineering and design leads Product: Huddle Status: Draft

1. Summary

Huddle will generate a post-meeting summary that highlights decisions and action items, identifies owners when it can do so reliably, and can email confirmed action items to their owners automatically. This replaces the manual note-taking people do after calls.

The core safety requirement is non-negotiable: an action must never be assigned to someone who did not attend the meeting. Calendar invitations alone cannot establish attendance: some invitees do not join, and around 15% of meetings include someone who was not invited. Speaker labels also cannot establish identity on their own; they are wrong about 8% of the time, especially during interruptions or when several people share a meeting-room connection.

The product must therefore separate three questions: what was decided, what needs to be done, and who—if anyone—can be safely identified as the owner. When attendance or ownership is uncertain, Huddle must leave the action unassigned rather than guess.

2. Problem and goals

Today, attendees write their own notes after meetings. This is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.

Goals - Produce a concise, useful summary after a recorded meeting. - Clearly distinguish decisions from discussion, proposals, and unresolved questions. - Extract action items and assign them only to verified meeting attendees. - When enabled, email confirmed action items to their owners shortly after the meeting ends. - Let users correct mistakes and give feedback.

Non-goals for the first release - Replacing full meeting transcripts or recordings. - Creating tasks in external project-management tools. - Inferring deadlines, priorities, or owners that were not stated. - Automatically assigning work when the owner is unclear. - Resolving disagreements about whether a decision was made without review.

3. Users and primary workflow

Users: meeting attendees who need an accurate recap; organizers who want to review or share it; and action owners who need a clear follow-up.

  1. A recorded meeting ends. Huddle gathers the transcript, speaker labels, calendar context, and actual session attendance data.
  2. Huddle generates a summary with separate Decisions, Action items, and, where useful, Open questions sections.
  3. For each action, Huddle identifies the supporting transcript evidence and checks the proposed owner against the verified attendee roster.
  4. If automatic action-item email is enabled for the organization, Huddle emails only actions with a sufficiently clear owner whose attendance is verified.
  5. The organizer and attendees can view the summary, correct it, and report errors. Unassigned actions remain visible for follow-up but are not emailed as assigned work.

4. Functional requirements

4.1 Establishing who attended

  • Use Huddle’s meeting-session join/leave records as the attendance source of truth. The calendar invite is context, not proof that someone attended.
  • Include identifiable people who joined without a calendar invitation. Exclude invitees who did not join.
  • Represent an attendee using a verified Huddle account, authenticated guest identity, or another approved identity mechanism.
  • If a meeting-room device represents several people and Huddle cannot reliably map speakers to individuals, do not treat the device name or a guessed speaker identity as a person. The organizer may confirm identities or owners in the summary UI.
  • If reliable session attendance data is unavailable, fail closed: generate the summary, but do not assign or email action items.

4.2 Generating decisions and action items

  • Generate a short overview and the three optional sections: Decisions, Action items, and Open questions.
  • Include a decision only when the transcript supports an explicit agreement or choice. Label proposals or unresolved discussion as open questions, not decisions.
  • Phrase each action as a concrete task. Include a due date only if one was explicitly stated; otherwise show no due date.
  • Attach a brief transcript excerpt and timestamp to each decision and action so users can check the source.
  • Do not invent missing details. If the task is clear but no owner is, show it as Owner not confirmed. If the task itself is ambiguous, flag it for review rather than present it as certain.

4.3 Assigning owners safely

  • Every proposed owner must match a verified person in the actual attendee roster. Enforce this as a backend validation rule, not just a model instruction or UI check.
  • A calendar invitee who did not join must never be assigned. A person who joined without an invite may be assigned if their identity and attendance are verified.
  • Speaker labels may help identify who volunteered, but must not independently authorize an assignment. When overlap, room audio, or attribution uncertainty makes ownership unclear, leave the action unassigned.
  • Users may assign or change an owner in the UI only to a verified attendee. Record the change and its actor in the audit log.
  • Recheck the attendance invariant before saving an assignment and immediately before sending an email.

4.4 Automatic emails

  • Provide an organization-level setting to enable automatic action-item emails. Once enabled, send an email after generation to each verified owner with at least one confirmed action.
  • Do not send an assignment email for an unassigned action or to anyone not verified as having attended. Do not send the full transcript by default.
  • Include the meeting title and date, the recipient’s confirmed actions, any explicitly stated due dates, a link to the Huddle summary, and a way to report an incorrect assignment.
  • If the meeting has no verified owners, send no owner emails. The organizer may receive a link to the summary according to organization settings.
  • Target delivery within five minutes of the meeting ending at the 95th percentile. If generation or identity checks fail, do not send speculative emails; show the organizer that the summary or email is delayed or needs review.
  • Provide clear controls to turn automatic emails off. Email delivery status and failures must be visible to the organizer.

5. UX requirements

The summary should make confidence and uncertainty understandable without exposing raw model scores. Show an Owner confirmed state for verified assignments and Owner not confirmed when a person cannot be safely identified. Present source timestamps as links into the recording or transcript, subject to existing access permissions.

Organizers must be able to edit summary text, remove an incorrect item, or assign an unassigned item to a verified attendee. Edits must be saved with a clear indication of who made them. If an action email has already been sent, do not silently replace the original assignment: notify the organizer that the owner changed and offer a corrected email to the new verified owner. Never email the former owner again unless explicitly requested.

The UI should explain that the calendar list may include people who did not attend and that unidentified room participants require confirmation. Use clear empty states when there are no decisions or actions.

6. Data, privacy, and operational requirements

Use only meeting data already available to Huddle for recording, transcription, and attendance, subject to existing recording consent and organization policies. Apply existing access controls to summaries, transcript evidence, and recording links. Store generated summaries, edits, assignment records, email status, and relevant audit events under Huddle’s existing retention and deletion policies. Do not use action-item content to train models unless the organization’s existing data policy explicitly permits it.

Generation and email must be idempotent: retries must not create duplicate summary items or send duplicate emails. Log the evidence reference, proposed owner, roster-validation result, any human correction, and delivery outcome for troubleshooting and audit.

7. Acceptance criteria

The release is acceptable only if: 1. A calendar invitee who did not join is never assigned or emailed an action. 2. An identifiable, authenticated attendee who joined without an invite can be assigned when the transcript supports the assignment. 3. An unclear speaker in overlapping audio produces an unassigned action, not a guessed owner. 4. A shared-room connection with no reliable person mapping cannot result in an individual assignment without organizer confirmation. 5. Every emailed action has a verified attendee owner and a link to its supporting summary. 6. Email retries do not result in duplicate messages. 7. Users can correct an assignment, and the correction is restricted to verified attendees.

8. Success metrics and rollout

Track summary generation success, time to summary, owner-confirmation rate, action-item correction and removal rates, email delivery time, duplicate-email rate, and user feedback. The primary safety metric is the number of assignments or assignment emails to non-attendees; the target is zero. Also monitor missed or unassigned owners to understand the trade-off between safety and usefulness.

Roll out behind an organization-level feature flag. Start with internal meetings, then a small customer cohort. Review false assignments and user corrections before expanding. Do not broaden automatic-email availability until attendance validation, idempotency, and failure handling pass the acceptance criteria.

9. Dependency to resolve before implementation

Engineering must confirm that Huddle can provide reliable, person-level join/leave records, including authenticated guests and shared-room behavior. If that data is unavailable or insufficient for a meeting, the safe behavior is to generate the summary without assigning or emailing owners.

Claims the judge couldn’t find in the brief

  • Some invitees do not join meetings.
  • Huddle can provide reliable, person-level join/leave records as the attendance source of truth.
  • Today’s manual note-taking is inconsistent, takes time, and makes it easy for decisions and follow-ups to be lost.
  • Huddle has existing meeting-session join/leave records and actual session attendance data.

Grades and run details

Decision model 72 · LLM judge 2 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly11%
  • passAddresses the actual decision90%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty86%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable78%
  • passSpecifies AI failure behaviour100%
  • partialSuccess settles trade-offs88%
  • partialProposes tests that could fail68%
Run
Run
#1
API response time
34 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Respects explicit constraintsRightMixedMixed
Sonnet 5.5 · API

It respects the non-attendee constraint through roster validation, deterministic checks, UI picker restrictions, and the absence of auto-emailing, and the PRD is the right form and length for the named readers.

Gemini 3.5 Flash-Lite · Gemini

The constraint that summaries never assign actions to non-attendees is enforced via verification and fallback to unassigned/host; the PRD also fits the requested form and length.

GPT-6 Luna · API

It violates the hidden prohibition on auto-emailing action items without human review and likely exceeds the requested 1,500-word length.

Identifies material uncertaintyRightWrongMixed
Sonnet 5.5 · API

It names key unknowns such as presence-data quality, label-error thresholds, room attestation, and guest privacy, and says how they would be resolved or what would change the call.

Gemini 3.5 Flash-Lite · Gemini

It mentions error rates but does not specify what conditions would change the design decision (e.g., an override-rate threshold that would trigger disabling auto-email).

GPT-6 Luna · API

It names speaker-label error and missing attendance data but does not specify attribution confidence thresholds or how those uncertainties would be resolved quantitatively.

Produces the required deliverableRightRightMixed
Sonnet 5.5 · API

A complete PRD for engineering and design leads is present with problem, requirements, UX, metrics, and milestones, and it is actionable as-is.

Gemini 3.5 Flash-Lite · Gemini

The output is a structured PRD with functional specs, UI mockups, and technical architecture, within 1,000–1,500 words, and is usable by engineering and design leads.

GPT-6 Luna · API

The PRD is complete in form but exceeds the requested length and depends on attendance data that is not confirmed to exist.

Success settles trade-offsRightWrongWrong
Sonnet 5.5 · API

It provides measurable targets with time frames, including zero non-attendee assignments and 5-minute delivery, and explicitly makes non-assignment a hard constraint over coverage.

Gemini 3.5 Flash-Lite · Gemini

No explicit trade-off rule is given (e.g., accepting lower coverage to keep assignment precision above a stated level), and not all success metrics have targets and time frames.

GPT-6 Luna · API

It gives a safety target and delivery target but lacks a time frame for success metrics and does not explicitly state which trade-off rule wins beyond vague monitoring.

All got wrong 3

Uses the supplied evidence correctlyWrongWrongWrong
Sonnet 5.5 · API

It relies on unsupported claims about session join/leave logs, note quality, and room-device behavior that are not in the supplied context.

Gemini 3.5 Flash-Lite · Gemini

The output claims Huddle is used by 'enterprise and mid-market companies' and has a 'Huddle Web & Desktop App', neither of which is in the supplied context.

GPT-6 Luna · API

It relies on unsupported current-situation claims such as reliable join/leave records and no-shows, which are not in the supplied context.

Avoids unsupported claimsWrongWrongWrong
Sonnet 5.5 · API

It presents interpretations and current-state facts such as late/inconsistent notes and session logs as established when the evidence only says people write their own notes and does not mention logs.

Gemini 3.5 Flash-Lite · Gemini

It presents 'enterprise and mid-market' and the existence of a desktop client as established facts without backing from the supplied evidence.

GPT-6 Luna · API

It presents unsupported claims about no-shows, join/leave records, and note-taking quality as established facts.

Proposes tests that could failWrongWrongWrong
Sonnet 5.5 · API

Some gates lack numeric thresholds or explicit actions, especially the auto-send 'agreed bar' and post-launch metrics, so not every proposed test can fail cleanly.

Gemini 3.5 Flash-Lite · Gemini

No tests, gates, or kill criteria with numeric thresholds, measurement windows, and consequents are proposed.

GPT-6 Luna · API

Acceptance criteria are mostly binary and lack numeric thresholds, measurement windows, or clear kill/pilot actions.

All got right 2

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

It commits clearly to host review before sending, explicitly rejects fully automatic action-item emails for v1, and states the conditions that would allow later automation.

Gemini 3.5 Flash-Lite · Gemini

The PRD commits to a clear design, addressing the Sales request with an opt-in auto-email and a review window, and is unambiguous for engineering and design leads.

GPT-6 Luna · API

It commits early to safe owner verification with optional auto-email and states conditions based on join/leave data availability.

Specifies AI failure behaviourRightRightRight
Sonnet 5.5 · API

It specifies unassigned defaults, confidence tiers, validator rejections, low-confidence suggestions, host review, and an owner 'This isn't mine' correction path.

Gemini 3.5 Flash-Lite · Gemini

It defines UI warning states for unverified owners, low confidence, and fallback to raw transcript on LLM timeout.

GPT-6 Luna · API

It specifies fail-closed behavior, unassigned actions when uncertain, and no speculative emails.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD