Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.9 · synthetic · AI product, support

Needs more than a quick edit8 of 10 checks passed

Unsupported claims about current workflow and data content would require re-checking evidence, and the data section's assumptions need validation before engineering can act.

Output

PRD: AI Draft & Route for Support Tickets

Sep 24, 2026 · @Martin Slaney

1. Summary and problem

Ledgerly's support team (42 agents, four queues: Billing, Technical, Account access, General) handles \~9,000 tickets a week. Average first response time (FRT) is 7 hours against a target of under 2.

Most of that delay is not writing time. It is triage and lookup1: \~60% of tickets are answered with one of \~40 saved replies, and 38% are billing. Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details.

This feature does two things when a ticket arrives:

  1. Routes it to the correct queue, with a confidence score.
  2. Drafts a first response, grounded in the saved-reply library, the customer's account data and similar past resolutions, for an agent to review.

Nothing is sent without an agent approving it. On top of that, any draft containing a refund commitment is hard-gated per Legal (§4.3).

2. Goals, non-goals and success metrics

Goals

  • Bring average FRT under 2 hours without adding headcount.
  • Cut misrouted tickets and the reassignment delay they cause.
  • Reduce agent handling time on saved-reply tickets, so agents spend more time on the \~40% that need real investigation.

Non-goals (v1)

  • Auto-sending any response, of any type.
  • Drafting follow-ups beyond the first response.
  • Taking actions in Ledgerly (issuing refunds, resetting passwords, changing plans). The model drafts text; agents act.
  • Customer-facing chatbot or deflection.

Success metrics (measured per queue, against a 4-week pre-launch baseline)

MetricTargetGuardrail
Average FRT< 2hP90 FRT must not rise
Routing accuracy (final queue = predicted queue)≥ 92% overall, ≥ 90% per queueAccount access recall ≥ 95%
Draft acceptance (sent with light or no edits)≥ 50% of drafted tickets—
Agent handling time on drafted tickets−30%—
CSAT on drafted ticketsNo drop vs baseline (±1pt)Reopen rate not up >1pt
Refund commitments sent without refund-approval step0Hard requirement

"Light edit" is defined as a normalised edit distance below 0.2 between draft and sent text. ML owns the metric definition; Support Ops signs it off before pilot.

3. Users and core workflow

Users: support agents (reviewers of every draft); queue leads (monitor routing, handle overrides); Support Ops (owns saved replies, policies and the refund-approval rota).

Flow for a new ticket

  1. Ticket created (email or in-app form).
  2. Router predicts queue + confidence within 30s. High confidence → assigned to that queue. Low confidence → Triage view for a lead to assign in one click.
  3. Drafter produces a first response and attaches it to the ticket as an internal draft, with: the saved reply(s) it drew on, account facts it inserted, and any flags (refund, low confidence, missing data).
  4. Agent opens the ticket, reviews the draft, and chooses Send, Edit & send, or Discard (with a reason code).
  5. If the draft is refund-flagged, Send is replaced by Request refund approval (§4.3).
  6. Final queue, sent text and action are logged as training and evaluation signal.

4. Functional requirements

4.1 Routing

  • R1. Classify every new ticket into Billing, Technical, Account access or General, with a calibrated confidence score.
  • R2. Auto-assign when confidence ≥ threshold (set per queue from eval, starting target: ≥ 95% precision at that threshold). Below threshold → Triage view.
  • R3. Account access is the costliest miss (locked-out customers).3 Tune for recall on this class; a ticket with any access signal and ambiguous classification goes to Account access, not General.
  • R4. Agents can reassign in one click; every reassignment is logged with the original prediction.
  • R5. Support Ops can switch routing to suggest-only per queue without a deploy.

4.2 Drafting

  • D1. Generate a draft for every routed ticket in English within 60s of creation. Other languages: no draft in v1, flag only.
  • D2. Retrieval first: identify the best-matching saved reply (or "none"). When one matches, the draft is that reply personalised with ticket and account context, not free text. When none matches, draft from similar resolved tickets and help-centre articles, and label it Free-form draft.
  • D3. Account facts (plan, billing dates, invoice amounts, last payment status) come only from read-only lookups against Ledgerly's billing/account APIs, never from model memory. Any fact the model could not verify is left as a visible `[placeholder]` that blocks sending until filled.
  • D4. Show sources inline: saved-reply ID, linked tickets/articles, API fields used.
  • D5. Never promise timelines, credits, discounts, policy exceptions or refunds unless the saved reply itself contains them2 (refunds additionally gated, §4.3).
  • D6. Skip drafting (flag only) for: legal threats, suspected fraud, data deletion/GDPR requests, security incidents, and abusive or distressed customers. Keyword + classifier; list owned by Support Ops.

4.3 Refund guardrail (Legal requirement)

Legal requires no automated sending of refund commitments. Because every draft already needs agent approval, we go further so a refund promise cannot slip through in a routine approve click:

  • G1. A dedicated refund-commitment detector runs on every draft and on the final edited text at send time (agents may add refund language themselves). Tuned for recall ≥ 99% on a Legal-reviewed test set; false positives are acceptable.
  • G2. If triggered, the ticket cannot be sent via one-click approval. The agent must confirm the refund is authorised under current policy (Support Ops defines who can authorise which amounts) via a separate confirmation step that is logged.
  • G3. The drafter must never generate a refund commitment from free-form reasoning; refund language may only come from approved refund saved replies.
  • G4. No bulk-approve action exists anywhere in the product.
  • G5. Legal reviews the detector test set and the confirmation UX before pilot, and receives a monthly log of refund-flagged sends.

Open for Legal: does "automated sending" cover a one-click approve by an agent? This PRD assumes one-click approval is not automated but adds G2 as defence in depth. Confirm before build.

4.4 Agent experience

  • A1. Draft appears pre-filled in the reply box, visibly marked as AI-drafted until edited.
  • A2. Discard requires a reason: wrong answer, wrong tone, missing info, wrong queue, should not be drafted.
  • A3. Agents never lose their normal tools; saved replies remain available manually.
  • A4. Customers are not told a draft was AI-assisted (agent authors the sent message). Support Ops to confirm this against Ledgerly's AI disclosure policy.

5. ML approach, data and evaluation

Data. Two years of resolved tickets (\~900k) with final queue, agent replies and saved-reply usage.

  • Label routing from the final queue, not the initial one; tickets that were reassigned are the most valuable examples.
  • Map historic replies to saved-reply IDs where possible (exact/near-match), giving a supervised "which reply fits" dataset.
  • Down-weight or exclude replies older than any policy or pricing change; Support Ops supplies the change dates. Stale answers are the main quality risk.
  • Strip PII before any use outside the production data boundary; confirm data-processing terms with Legal/DPO for customer ticket content.

Approach. Start with the simplest thing that hits the targets: an LLM classifier (or fine-tuned small model if cost/latency requires) for routing; retrieval over saved replies + resolved tickets + help centre, then LLM personalisation for drafting; a separate refund detector. Model choice is ML's call, constrained by the latency targets in §3 and D1 and per-ticket cost.

Offline evaluation (gate to pilot)

  • Held-out set: most recent 3 months, time-split (no leakage).
  • Routing: per-queue precision/recall and calibration; auto-assign coverage at the chosen threshold.
  • Drafting: saved-reply selection top-1 accuracy; 500-ticket human-graded sample scored by senior agents (correct, complete, on-policy, tone) with a ≥ 80% "send with light or no edit" bar.
  • Refund detector: recall ≥ 99% on Legal-reviewed set.
  • Factual accuracy: 0 invented account facts in the graded sample.

Online monitoring. Dashboards per queue for acceptance, edit distance, discard reasons, reassignments, FRT, CSAT and refund flags. Weekly review of a random sample of sent drafted replies by Support Ops QA.

6. Rollout, risks and open questions

Rollout

PhaseScopeExit criteria
0. Shadow (2 wks)Route + draft on all tickets, invisible to agentsOffline bars met on live traffic
1. Pilot (3 wks)Billing queue, 8 agents; routing suggest-onlyAcceptance ≥ 40%, no CSAT drop, 0 refund-guardrail misses
2. ExpandAll queues drafting; routing auto-assign above thresholdMetrics in §2 hold for 2 consecutive weeks
3. Steady stateMonthly threshold review, quarterly saved-reply refresh—

Billing goes first: highest volume, most saved-reply coverage, and it exercises the refund guardrail early under close watch.

Risks

  • Rubber-stamping. Agents approve drafts without reading them. Mitigate: placeholders block send, refund gate, QA sampling, track time-on-draft before send.
  • Stale policy in drafts. Mitigate: saved replies as primary source, policy-change exclusion in training data, owner for saved-reply freshness.
  • Misrouting Account access. Mitigate: recall-weighted threshold (R3), Triage fallback.
  • Agent trust. Early bad drafts kill adoption. Mitigate: shadow phase, drafts only above a quality-confidence score; otherwise show the suggested saved reply only.

Open questions

  1. Legal: does agent one-click approval count as "automated sending" (§4.3)?
  2. Support Ops: who can authorise refunds, and at what amounts?
  3. Eng: which helpdesk platform APIs support pre-filled internal drafts and blocking send?
  4. DPO: approval to use historic ticket content with a model provider.
  5. Support Ops: AI-assistance disclosure to customers (A4).

Owners: Eng lead (integration, guardrail enforcement, UI), ML lead (models, eval, monitoring), Support Ops lead (saved replies, policy, pilot, QA).

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisSubstantial rework

    Source: VolumesAverage first response is 7 hours against a target of under 2.

    What we’d changePresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: RiskLegal requires no automated sending of refund commitments.

    What we’d changeState the refund rule once and consistently: generated drafts never contain refund commitments, and any commitment an agent adds goes through the separate confirmation step.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changePresent it as an assumption to check with Support Ops. The brief doesn't rank the cost of misrouting by queue, and the routing design leans on it.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyThe PRD states that the delay cause is triage and lookup and that tickets wait unassigned, and assumes saved-reply usage is in the historical data, none of which is evidence from the brief.
  • Avoids unsupported claimsPresents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
Got right · 8
  • Addresses the actual decisionCommits to building the feature with phased rollout and exit criteria that would halt expansion if not met, framed for engineering/ML/support-ops leads.
  • Respects explicit constraintsRespects all constraints: addresses the intended readers, within word count, mandates agent approval, and enforces no automated refund commitments via guardrails.
  • Identifies material uncertaintyIdentifies open questions about legal definition, refund authorizations, platform APIs, data use, and AI disclosure, with owners and action to resolve.
  • Produces the required deliverableProvides a complete PRD for the required audience, within the 1,000–1,500 word range, that they could act on.
  • Specifies AI failure behaviourDefines low-confidence routing to a Triage view, placeholders blocking send, draft suppression for sensitive cases, and discard reasons.
  • Defines evals and a launch barSpecifies offline evaluation with human-graded sample and 80% send-with-light-edits bar, routing precision/recall thresholds, and a refund-detector recall ≥99%.
  • Success settles trade-offsSets success metrics with targets (FRT <2h, routing accuracy ≥92%) and explicit trade-off rules like prioritizing account access recall at the expense of other queues.
  • Proposes tests that could failEach pilot phase has numeric exit criteria (e.g., acceptance ≥40%, 0 refund misses), measurement windows (3 weeks), and triggers that stop rollout if not met.

Claims the judge couldn’t find in the brief

  • Most of that delay is not writing time. It is triage and lookup
  • Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details
  • Two years of resolved tickets (~900k) with final queue, agent replies and saved-reply usage are available
  • Account access is the costliest miss (locked-out customers)
  • Billing queue has the most saved-reply coverage

Grades and run details

Decision model 75 · LLM judge 8 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly46%
  • passAddresses the actual decision69%
  • passRespects explicit constraints35%
  • passIdentifies material uncertainty87%
  • failAvoids unsupported claims71%
  • passProduces the required deliverable69%
  • passSpecifies AI failure behaviour97%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs80%
  • partialProposes tests that could fail71%
Artefacts
Run
Run
#1
Time to output
2.0 min
Submitted
24 Sept 2026

Needs more than a quick edit4 of 10 checks passed

Contains an unsupported current routing accuracy figure, missing evaluation plan/launch bar, and no trade‑off rules; a PM would need to re‑check evidence and add substantial missing sections.

Output

Product Requirements Document (PRD): AI-Powered Ticket Triage and Response Draft Generation ("AutoDraft & Route")

---

1. Document Overview

MetadataDetails
:---:---
Project NameAutoDraft & Route (Ledgerly Support Intelligence)
Target AudienceEngineering, Machine Learning (ML), and Support Operations (Ops) Leads
Document StatusDraft / Ready for Engineering Review
Target LaunchQ1 2027
GoalReduce Average First Response Time (FRT) from 7 hours to <2 hours while maintaining human-in-the-loop safety and strict legal compliance on financial commitments.

---

2. Background & Problem Statement

2.1 Current State & Metrics

Ledgerly is a premier accounting software platform designed specifically for small businesses. As our customer base scales, our support infrastructure is facing severe volume pressures: * Support Team Size: 42 agents distributed across four specialized queues: Billing, Technical, Account Access, and General. * Ticket Volume: ~9,000 incoming tickets per week, with 38% (~3,420 tickets/week) falling into the Billing queue. * First Response Time (FRT): Current average is 7 hours, severely lagging behind our target of under 2 hours. * Repetitive Workflows: Approximately 60% of all incoming tickets are resolved using one of roughly 40 standard saved replies (e.g., password resets, invoice retrieval, subscription tier changes).

2.2 The Opportunity

We possess two years of historical resolved tickets, including full conversation histories and agent replies. This corpus serves as a robust training and evaluation dataset for an ML-driven system that can: 1. Accurately classify and route incoming tickets to the correct queue. 2. Automatically draft context-aware first responses using our proven saved replies and historical resolutions.

2.3 Core Constraint (Legal Compliance)

  • Zero Automated Sending: Legal has mandated a strict human-in-the-loop requirement. The AI system must never autonomously send messages to customers.
  • Financial Safety: The model is strictly prohibited from making automated commitments regarding refunds, credits, or billing adjustments without explicit human review and approval.

---

3. Goals and Success Metrics (KPIs)

MetricCurrent BaselineTarget (3 Months Post-Launch)
:---:---:---
Average First Response Time (FRT)7 hours< 2 hours
Routing Accuracy~75% (Manual/Rule-based)1> 95% across all four queues
Draft Acceptance RateN/A> 70% (used as-is or with minor edits)
Agent Handle Time (AHT)Baseline-25% reduction in time spent crafting initial replies
Compliance / Unauthorized Sends00 violations (Hard safety guardrail)

---

4. User Personas & Workflows

4.1 Support Agent

  • Role: Resolves customer queries, reviews AI drafts, edits when necessary, approves responses, and handles complex escalations.
  • Workflow: Opens the ticketing dashboard, sees incoming tickets pre-sorted into their designated queue with an AI-generated draft response attached. The agent reviews the draft, clicks "Approve & Send" or modifies the text, ensuring zero financial compliance risk.

4.2 Support Operations Lead

  • Role: Manages queue health, tracks SLAs, audits agent performance, and maintains the library of saved replies and routing rules.
  • Workflow: Monitors routing precision dashboards, updates macro/saved reply templates, and analyzes draft acceptance/rejection feedback loops to refine prompts.

4.3 Engineering & ML Leads

  • Role: Build, deploy, scale, and monitor the inference pipeline, classification models, LLM orchestration, and UI integration.

---

5. Functional Requirements

5.1 Automated Ticket Routing (Triage Engine)

  • Queue Classification: Upon ticket ingestion, the ML classification service must categorize incoming text into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • Confidence Thresholds:
  • Confidence $\ge$ 85%: Automatically assign the ticket to the predicted queue.
  • Confidence < 85%: Route the ticket to a "Triage Review" sub-queue or default General queue with a visual tag flagging low classification confidence.
  • Metadata Tagging: Tag tickets with sub-intent labels (e.g., `Billing: Refund Request`, `Technical: Bank Sync Failure`) to assist the response generation engine.

5.2 Response Draft Generation

  • Context Window & Retrieval: The generation engine must leverage the incoming user query, customer account metadata (plan tier, active status), and relevant historical examples from the 2-year resolved ticket dataset (via RAG / vector search).
  • Saved Reply Integration: The model must prioritize mapping inquiries to Ledgerly's 40 core saved replies where applicable, maintaining brand voice, accuracy, and tone consistency.
  • Draft Status: Every generated response must be marked with a distinct internal status: `Draft - Pending Agent Review`. It must remain locked from customer view until explicit human sign-off.

5.3 Agent Workspace UI/UX Integration

  • Side-by-Side Review: The agent UI must display the incoming customer message alongside the AI-generated draft in a clear, editable text box.
  • Action Controls:
  • [Approve & Send]: Immediately transmits the draft to the customer and logs the interaction.
  • [Edit & Send]: Allows agents to modify text inline before sending. Edits must be logged for ML fine-tuning feedback loops.
  • [Discard & Rewrite]: Clears the draft if the AI misunderstood the prompt.
  • Visual Safety Badges: Clear UI warnings must appear on any ticket flagged as containing billing or refund keywords, reminding agents of compliance protocols.

---

6. Non-Functional & Safety Requirements

6.1 Legal & Financial Guardrails (Critical)

  • No Auto-Dispatch: Zero API endpoints or automation rules are permitted to transition a draft state to `Sent` without a cryptographic token or database action originating from an authenticated human agent session.
  • Refund Keyword Interceptor: A deterministic rule-based filter must scan all generated drafts for high-risk financial terminology (`refund`, `credit`, `waiver`, `chargeback`). If detected, the UI must highlight the section in amber and require a secondary manual checkbox confirmation from the agent before sending.

6.2 Performance & Latency

  • Inference Latency: End-to-end classification and draft generation must complete within < 3.0 seconds of ticket ingestion via webhook, ensuring drafts are ready before an agent opens the ticket.
  • Availability: 99.9% uptime SLA for the ML inference and routing service.

6.3 Data Privacy & Security

  • PII Redaction: Customer tax IDs, banking credentials, and sensitive financial figures must be handled in compliance with GDPR, CCPA, and Ledgerly security policies. Training data pipelines must scrub extraneous PII prior to embedding generation.

---

7. Data, ML, & Training Strategy

7.1 Dataset Utilization

  • Corpus: 2 years of historical resolved tickets (approx. 900,000+ interaction pairs across 42 agents).
  • Data Splitting: 80% training/fine-tuning, 10% validation, 10% held-out test set reflecting recent seasonal billing spikes.

7.2 Model Architecture & Approach

  • Classification: Fine-tuned encoder model (e.g., BERT/DeBERTa variant) for multi-class queue routing and intent tagging.
  • Generation: Instruction-tuned LLM accessed via secure API or private VPC deployment, augmented with a vector database (RAG) containing the 40 approved saved replies and top-performing historical resolutions.
  • Feedback Loop: Log all agent edits (diffs) and rejections to construct a weekly RLHF (Reinforcement Learning from Human Feedback) or supervised fine-tuning dataset to reduce hallucination rates.

---

8. High-Level Architecture & Data Flow

``` [Customer Email / Web Form] │ ▼ [Ingestion Service / Webhook] │ ├──► [ML Classification Engine] ──► Assigns Queue & Sub-intent │ └──► [RAG Vector DB (Saved Replies & History)] │ ▼ [LLM Draft Generation Service] ──► Generates Initial Response Text │ ▼ (Strictly Locked: Status = Pending Review) [Support Ticketing Database] │ ▼ [Agent UI Dashboard] ◄─── Human Agent Reviews, Edits, & Approves │ ▼ (Explicit Human Action) [Customer Dispatch API] ```

---

9. Phased Rollout Plan

  • Phase 1: Shadow Mode (Weeks 1–4)
  • Deploy classification and draft generation in the background.
  • Log routing accuracy and draft quality without surfacing drafts to agents. Establish baseline ML accuracy metrics.
  • Phase 2: Internal Pilot / Beta (Weeks 5–8)
  • Roll out the UI feature to a pilot group of 5 senior Billing agents.
  • Measure acceptance rates, time saved, and friction points. Refine prompt engineering and guardrails.
  • Phase 3: General Availability Across All Queues (Weeks 9–12)2
  • Roll out to all 42 agents across Billing, Technical, Account Access, and General queues.
  • Establish continuous monitoring dashboards for Support Ops and Engineering leads.

---

10. Open Questions & Risks

  1. Edge Cases in Billing: How should the model handle complex multi-invoice dispute threads where historical context spans multiple months? (Mitigation: Surface the last 3 ticket summaries alongside the draft).
  2. Agent Adoption: Will agents trust the AI drafts, or will they rewrite them from scratch? (Mitigation: Emphasize time savings during training and incorporate agent feedback buttons).

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimStart again

    Source: Volumes9,000 tickets a week; 38% are billing.

    What we’d changeRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.

  2. 2Test or gate too weakTighten the testSubstantial rework

    What we’d changeAdvance each phase on quality gates (routing precision, draft acceptance, zero refund-commitment misses) rather than the calendar, with a human grading rubric and a rollback trigger.

Check by check

Got wrong · 6
  • Uses the supplied evidence correctlyStates current routing accuracy of ~75% as fact without any support from the supplied evidence and does not label it as an assumption.
  • Identifies material uncertaintyOpen questions are listed but are not tied to decision‑changing thresholds or explicit resolution plans; the PRD does not state what would halt or alter the rollout.
  • Avoids unsupported claimsThe routing accuracy baseline of ~75% is presented as a fact without being marked as an assumption, and no evidence supports it.
  • Defines evals and a launch barNo evaluation dataset criteria, metrics (e.g., draft quality), or specific launch threshold for progressing from shadow mode to pilot or GA are defined.
  • Success settles trade-offsSuccess metrics include targets and a timeframe, but no explicit trade‑off rule (e.g., coverage vs. precision) is stated.
  • Proposes tests that could failThe phased rollout lacks numeric thresholds, measurement windows, and actions tied to results; it does not define kill criteria.
Got right · 4
  • Addresses the actual decisionThe PRD commits to a clear design for the AI feature, framed for the engineering, ML, and ops leads.
  • Respects explicit constraintsThe document remains within the 1,000–1,500‑word range, addresses the required readers, and enforces legal constraints with human‑only dispatch and a refund‑keyword interceptor.
  • Produces the required deliverableThe output is a complete PRD with functional requirements, architecture, and rollout plan that the named leads could start work from.
  • Specifies AI failure behaviourLow-confidence classification routes to a triage queue, and a deterministic refund‑keyword interceptor triggers a secondary manual checkbox.

Claims the judge couldn’t find in the brief

  • Current routing accuracy is ~75% (Manual/Rule-based).

Grades and run details

Decision model 50 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly58%
  • passAddresses the actual decision51%
  • partialRespects explicit constraints40%
  • partialIdentifies material uncertainty90%
  • partialAvoids unsupported claims48%
  • partialProduces the required deliverable46%
  • partialSpecifies AI failure behaviour27%
  • partialDefines evals and a launch bar83%
  • partialSuccess settles trade-offs73%
  • partialProposes tests that could fail69%
Artefacts
Run
Run
#1
Time to output
14 s
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD