Needs more than a quick edit6 of 9 checks passed
The PRD is mostly usable but assumes session join/leave logs exist and repeats unsourced claims about note quality and room devices, so the evidence would need re-checking. Some success criteria also lack exact thresholds and actions.
Output
PRD: AI Meeting Summaries with Decisions and Action Items
Status: Draft for engineering and design leads Product: Huddle
1. Problem
Huddle records and transcribes meetings with speaker labels, but attendees still write their own notes afterward. Notes are late, inconsistent, or never written, so decisions get lost and action items have no clear owner. We will generate a summary after each meeting that lists decisions and proposes action items with owners.
2. Goals and non-goals
Goals - Produce a summary within 5 minutes of meeting end: decisions, action items (task, owner, due date if stated), and open questions. - Never assign an action item to someone who was not in the meeting (hard constraint, see §4). - Keep the host's review to under 60 seconds for a typical meeting.
Non-goals (v1) - Fully automatic emailing of action items to owners (see §3). - Cross-meeting task tracking or integrations with Jira, Asana, and similar tools. - Summaries for meetings without a transcript.
3. Decision on the Sales request
Sales asked for action items to be emailed automatically to owners the moment a meeting ends. We are not building that for v1. Instead, we will send the draft summary to the host at meeting end. The host can send it to owners with one click, and owners are emailed on send.
Reasons: - Speaker labels are wrong about 8% of the time. An automatic email would sometimes tell a real person they own a task they never agreed to. Because it goes out with no human check, the error becomes visible to a customer's colleagues. - Email is irreversible, so a wrong or non-attendee assignment cannot be quietly fixed afterward. - The underlying Sales need is that owners hear about their tasks quickly and nothing falls through the cracks. A host prompt at meeting end with one-click send meets that need at the cost of a few seconds.
Path to automation: Auto-send will be an opt-in, per-workspace setting, considered after launch once the gates in §9 are met. It would send only high-confidence, presence-verified items and leave everything else for host review. Engineering should build the send pipeline so that auto-send is later a policy change rather than a rebuild.
4. The hard constraint: who can own an action item
Rule: An action item's owner must be a participant who was actually in the meeting when the item was discussed. Anyone else is ineligible.
This rule cannot rely on the calendar invite, and it cannot rely on the model behaving well.
Attendance source. The eligible-owner roster is built from Huddle's session join/leave logs, not from the calendar invite. This handles both failure modes: - Invitees who never joined are not on the roster. - The roughly 15% of meetings with uninvited joiners include those people, matched to a user account when possible.
Presence at time of item. A person is eligible for an item only if they were in the session at the timestamp of the relevant utterance. Someone who left at minute 10 cannot own an item from minute 40.
Enforcement is structural, not prompt-based. 1. The extraction step receives the roster as a list of participant IDs and must return an owner ID from that list or `null`. 2. A deterministic validator runs after the model. It rejects any owner ID not on the roster or not present at the item's timestamp, and it converts that item to unassigned. 3. The UI and email layers accept only validated owner IDs. There is no free-text owner field on the model output path. 4. Names mentioned in the transcript who were not present ("have Dana look at it") produce an unassigned item. The text may keep the speaker's wording, but Dana gets no assignment and no email.
Shared rooms. A room device usually appears as one participant, so individuals in the room are not in the presence log. We cannot verify that a named person is in the room. Such items become unassigned by default, and the host can attest who was in the room during review (§6). That attestation counts as human-verified presence.
Uninvited joiners without an account (guests) can be eligible owners, shown by display name. They receive email only if the host supplies an address during review. We never guess addresses.
5. Handling speaker-label errors
The 8% label error rate is concentrated in crosstalk and shared rooms. Eligible-owner enforcement does not solve it, because a mislabel can still pick the wrong person who was present. Mitigations:
| Signal | Handling |
|---|---|
| Owner named explicitly ("Priya, can you send the deck?") and Priya is present, followed by assent or no objection | High confidence. Name-based assignment does not depend on speaker labels. |
| Self-assignment ("I'll send it") | Depends on the label. High only if the segment is not in a crosstalk or shared-room span and label confidence is above threshold. Otherwise Medium. |
| Assignment inferred from context, or the segment overlaps other speakers | Low. Default to unassigned with a suggested owner the host can accept. |
| Speaker is a shared-room device | Never treated as a person. Unassigned unless a named individual is explicitly addressed and verified. |
Confidence tiers (High/Medium/Low) are stored on each item, shown to the host in review, and used later to gate auto-send. Thresholds are set empirically during the evaluation in §9. The tiers themselves are a v1 requirement.
6. User experience
Host flow 1. The meeting ends and the host gets a notification: "Summary ready to review." 2. The review screen lists: - Decisions as short statements, each linked to the transcript moment. - Action items with the task, owner chip, due date if stated, and a confidence indicator. - Unassigned items grouped at the top of the actions list, so gaps are visible. 3. The host can edit text, reassign from a picker limited to eligible people, delete items, or add room attendees ("Who was in the room?"). Adding room attendees makes them eligible for items in the room segments. 4. Send emails each owner their own items plus the decisions. Unassigned items go to no one.
Design requirements - Every item has a "jump to source" link that plays the transcript segment. Reviewing should be verification, not re-listening. - The owner picker never offers non-attendees. If the host wants to assign to an absent person, show a separate option, "Send as a follow-up (not from the meeting)," which is a different item type with clearly different wording and no "you agreed to" language. Design to decide whether this ships in v1; the default is that it does not. - Low-confidence owners appear as suggestions ("Suggested: Sam") and never as assignments. - Email copy says "Action item from Meeting name, date", includes the source excerpt, and offers a "This isn't mine" link that notifies the host. This gives owners a correction path. - If the host does not review, send a reminder at 24 hours. Nothing is sent to owners without host action in v1.
7. Functional requirements
- Roster service: builds the per-meeting roster with join/leave intervals from session logs and links to user accounts where possible.
- Extraction: from the transcript and roster, produce decisions, action items, and unassigned items. Each action item has task text, owner ID or null, due date or null, source timestamps, and a confidence tier.
- Validator: enforces §4 as code with unit tests, independent of the model, and logs every rejection.
- Review UI: as described in §6, on web and desktop. Mobile is view-only for v1.
- Delivery: per-owner emails on host send, with a record of who was emailed. Sends are idempotent, and edits after sending trigger an explicit "send update" action.
- Audit log: stores for each item the model output, validator result, host edits, and send events.
- Admin setting: workspace-level toggle for summaries, plus retention aligned with existing transcript retention.
8. Edge cases to specify and test
- Invitee never joined but is mentioned as owner.
- Uninvited joiner takes an action; the guest has no account.
- Person leaves and rejoins (multiple intervals).
- Two people share a first name; ownership must resolve by roster, and ambiguity means unassigned.
- Whole meeting in a shared room (all items default unassigned until the host attests).
- Multiple languages, or a transcript with large gaps.
- Nobody commits to anything: show "No action items detected" rather than inventing some.
9. Success metrics and launch gates
Launch gates (must all pass) - Zero non-attendee assignments across a labeled evaluation set of at least 500 meetings, deliberately oversampled for uninvited joiners, shared rooms, no-show invitees, and crosstalk. Any production occurrence afterward is a Sev-1 incident. - Among High-confidence items, wrong-attendee assignment rate of 3% or less (target, to be validated against the evaluation set; propose a stricter figure if data supports it). - Median summary delivery within 5 minutes of meeting end.
Post-launch metrics - Share of meetings where the host sends the summary within 1 hour. - Host edit rate per item, and owner "This isn't mine" rate. - Share of action items left unassigned. Some is healthy, but a high rate means the model is too timid.
Auto-send criteria (future): two consecutive months of High-tier precision above the agreed bar and no non-attendee incidents. Sales and Support are consulted before enabling.
10. Risks and open questions
- Sales disagreement. Sales may see the host step as friction. We should share §3 with them early and offer a pilot of auto-send with a few design-partner accounts after the gates are met.
- Hosts skipping review. If sending rates are low, the fix is better reminders and faster review, not removing the check.
- Presence data quality. Do our session logs reliably capture join/leave for all clients and dial-in users? Engineering to confirm in week 1.
- Guest privacy. Can we email or display guest names without consent, given customer policies? Legal to advise.
- Room attestation. Is host attestation adequate for the constraint, or should room participants confirm individually? Recommend host attestation for v1.
- Follow-up items for absent people. Ship in v1 or defer? Design and Product to decide.
11. Milestones
- Weeks 1–2: roster service, presence data audit, evaluation set labeling begins.
- Weeks 3–5: extraction, validator, confidence tiers, offline evaluation.
- Weeks 4–7: review UI and email delivery, in parallel.
- Weeks 8–9: internal dogfood, gate review.
- Week 10: limited beta with design-partner customers, including Sales' accounts.
Check by check
Got wrong · 3
- Uses the supplied evidence correctlyIt relies on unsupported claims about session join/leave logs, note quality, and room-device behavior that are not in the supplied context.
- Avoids unsupported claimsIt presents interpretations and current-state facts such as late/inconsistent notes and session logs as established when the evidence only says people write their own notes and does not mention logs.
- Proposes tests that could failSome gates lack numeric thresholds or explicit actions, especially the auto-send 'agreed bar' and post-launch metrics, so not every proposed test can fail cleanly.
Got right · 6
- Addresses the actual decisionIt commits clearly to host review before sending, explicitly rejects fully automatic action-item emails for v1, and states the conditions that would allow later automation.
- Respects explicit constraintsIt respects the non-attendee constraint through roster validation, deterministic checks, UI picker restrictions, and the absence of auto-emailing, and the PRD is the right form and length for the named readers.
- Identifies material uncertaintyIt names key unknowns such as presence-data quality, label-error thresholds, room attestation, and guest privacy, and says how they would be resolved or what would change the call.
- Produces the required deliverableA complete PRD for engineering and design leads is present with problem, requirements, UX, metrics, and milestones, and it is actionable as-is.
- Specifies AI failure behaviourIt specifies unassigned defaults, confidence tiers, validator rejections, low-confidence suggestions, host review, and an owner 'This isn't mine' correction path.
- Success settles trade-offsIt provides measurable targets with time frames, including zero non-attendee assignments and 5-minute delivery, and explicitly makes non-assignment a hard constraint over coverage.
Claims the judge couldn’t find in the brief
- Huddle has session join/leave logs that can build a per-meeting roster.
- Notes are late, inconsistent, or never written.
- A room device usually appears as one participant, so individuals in the room are not in the presence log.
- Email is irreversible.
Grades and run details
Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
- failUses the supplied evidence correctly53%
- passAddresses the actual decision99%
- passRespects explicit constraints67%
- passIdentifies material uncertainty97%
- failAvoids unsupported claims23%
- passProduces the required deliverable80%
- passSpecifies AI failure behaviour100%
- passSuccess settles trade-offs66%
- partialProposes tests that could fail65%
Run
- Run
- #1
- API response time
- 49 s
- Submitted
- 29 Sept 2026