Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.9 · synthetic · AI product, support

Usable with a quick edit9 of 10 checks passed

The PRD is usable as-is with only small edits to quantify the pilot go/no-go thresholds; the evidence and reasoning are sound and would not need redoing.

Output

PRD: AI-Drafted First Responses and Ticket Routing

Status: Draft for build planning Audience: Engineering, ML, and Support Operations leads Product: Ledgerly Support

1. Problem

Ledgerly receives about 9,000 support tickets a week, handled by 42 agents across four queues: Billing, Technical, Account access, and General. Billing is 38% of volume (about 3,400 tickets a week). Average first response is 7 hours against a target of under 2.

Two things likely drive the delay, and we don't yet know their split:

  • Wait time. Tickets sit in a queue, or in the wrong queue, before anyone opens them.
  • Handle time. Agents spend time writing or assembling a first reply. About 60% of tickets are already answered with one of roughly 40 saved replies, so much of this work is repetitive.

This feature has two parts. It routes each new ticket to the right queue, and it drafts a first response for the agent to review, edit, and send. Nothing is ever sent without an agent's approval.

2. Goals and non-goals

Goals 1. Bring mean first response time under 2 hours within 90 days of full rollout. 2. Route at least 92% of tickets to the correct queue on first assignment. 3. Have agents send at least 50% of drafts with no or light edits. 4. Ensure no draft or sent message contains an unreviewed refund commitment (see §5).

Non-goals - Auto-sending any message, in any queue, at any confidence level. - Resolving tickets, or taking actions such as issuing refunds, changing plans, or resetting access. - Replacing follow-up replies after the first response (later phase). - Changing queue structure or staffing.

3. Users

  • Support agents (42): review, edit, approve, or discard drafts, and correct routing.
  • Support-ops leads: own saved replies, policy content, thresholds, and quality review.
  • Legal/compliance: consulted on the refund guardrail and audit logs.

4. Functional requirements

4.1 Routing

  • R1. On ticket creation, classify into one of the four queues and assign it automatically.
  • R2. Store a confidence score with each assignment. Below a threshold set by support ops (initially tuned so about 10% of tickets fall below it), assign to General with a "needs triage" tag and show the top two suggested queues.
  • R3. Agents can re-route in one click and give an optional reason. Every re-route is logged as a labeled correction.
  • R4. Also emit tags used by drafting: a topic label (for example, invoice question, failed payment, login lockout) and a refund-related flag.

4.2 Draft generation

  • R5. A draft must be ready in the agent's ticket view within 60 seconds of ticket creation, and never block the ticket from appearing. If drafting fails or times out, the ticket appears as it does today.
  • R6. Use two drafting paths:
  • Saved-reply path. If a ticket matches one of the ~40 saved replies above a similarity threshold, use that reply with variables filled in (name, plan, invoice number). The wording stays as approved by support ops.
  • Generated path. Otherwise, generate a draft grounded in retrieved help-center articles and similar past resolved tickets. Every factual claim about policy, pricing, or product behavior must come from retrieved content. If the content is insufficient, the draft asks the customer a clarifying question or says an agent is investigating, rather than guessing.
  • R7. Each draft shows which path produced it and its sources (saved reply ID or article links) so agents can verify quickly.
  • R8. Match the tone and format of the best-rated historical agent replies. Support ops maintains the style guide.

4.3 Agent workflow

  • R9. Drafts appear pre-filled in the reply composer with three actions: Send (after optional edits), Discard, and Regenerate. Discarding takes an optional reason from a short list.
  • R10. The system has no auto-send code path. Sending requires an authenticated agent action, and this is enforced in the sending service, not just the UI.
  • R11. Log for each ticket: draft text, final sent text, edit distance, agent ID, time from ticket open to send, and discard or re-route reasons.

4.4 Account-access safeguards

  • R12. Drafts for Account access tickets must never confirm whether an account exists, reveal account details, or state that access was restored or changed. They may give standard verification instructions from approved content only.

5. Refund guardrail (legal requirement)

Legal requires no automated sending of refund commitments. Because agents approve every draft, the primary control is R10. We add layered controls so the model never puts a commitment in front of an agent as if it were approved policy.

  • G1. Prompt and template rules. Drafts may acknowledge a refund request and say it is being reviewed. They must not promise, imply, or estimate a refund, credit, waiver, or timeline, and no saved reply used by this feature may contain one.
  • G2. Output classifier. Every draft passes a commitment detector (rules plus model) before display. If it fires, replace the offending sentence with a neutral placeholder, such as "[Agent: refund decision needed]", and tag the ticket.
  • G3. Send-time check. The detector also runs on the final edited text. If an agent's own text contains a refund commitment, Send requires an explicit confirmation ("This message commits to a refund"), which is logged. Agents may make commitments in their own words. The system may not.
  • G4. Audit. Retain drafts and final messages for the period Legal specifies. Support ops reviews a weekly sample of 100 refund-related tickets.
  • G5. Release bar. The detector must reach at least 98% recall on a labeled set of refund-commitment phrasings, including implicit ones like "we'll take care of that charge", before Billing goes live. Legal signs off on the test set.

6. Data and evaluation

Available: two years of resolved tickets with agent replies.

Data preparation (ML lead owns) - Use the queue where a ticket was resolved as the routing label, not where it first landed, since misroutes are the problem we're fixing. Also keep the initial queue to measure the historical misroute rate. - Remove PII before training or indexing. Get security review of the vendor and hosting setup if any external model is used. - Filter out replies that are outdated (old pricing, retired features, superseded policies). Support ops flags policy change dates so the pipeline can exclude earlier replies. - Mark replies that used a saved reply. This gives the saved-reply matching set and a baseline for the other 40%. - Have support ops label about 1,500 recent tickets for refund-commitment presence, routing, and draft quality. This set doubles as the gold evaluation set.

Evaluation - Split by time, not randomly: train on the earliest ~21 months, test on the most recent 3. Random splits will overstate performance because of seasonality and policy drift. - Routing: report accuracy and per-queue precision and recall, plus confusion between Billing and General, and Account access and Technical, which are likely weak spots. - Drafts: blind human review by senior agents on a 5-point rubric (accuracy, policy correctness, tone, completeness), with automatic policy-violation checks. Compare against the original agent reply. - Slice results by queue, topic, and saved-reply versus generated path.

7. Success metrics

MetricBaselineTarget
Mean first response time7h<2h at 90 days post-rollout (<4h at pilot)
Routing accuracy (first assignment)Measure in shadow≥92%
Drafts sent unedited or lightly editedn/a≥50% overall; ≥75% on saved-reply path
Draft discard raten/a<15%
Policy-incorrect drafts in weekly reviewn/a<2%
Refund commitments sent without confirmation flown/a0
Customer satisfaction on first-response ticketsCurrent CSATNo decline

First task: instrument the current split of first-response time into queue wait and agent handle time. Drafting mainly cuts handle time. If wait time dominates, we also need queue-ordering changes, such as surfacing the oldest tickets first, or the 2-hour target will be missed even with high draft quality. Throughput is about 43 tickets per agent per day, so handle time matters, but it is not necessarily the bottleneck.

8. Rollout

  1. Shadow mode (2–3 weeks). Run routing and drafting on live tickets without showing anything to agents. Compare with actual outcomes, set thresholds, and validate the refund detector.
  2. Pilot (3 weeks). Enable routing for all tickets. Show drafts to about 12 agents across General, Technical, and Account access.
  3. Expand. Enable drafts for all agents in those three queues once the pilot meets the quality targets.
  4. Billing. Enable only after G5 is met and Legal has signed off. Start with 6–8 agents, then expand.

Any queue can be switched off by a feature flag within minutes. Rollback triggers: any refund-commitment incident, a policy-incorrect rate above 5% in weekly review, or a CSAT drop of more than 3 points.

9. Risks and open questions

  • Automation bias. Agents may approve drafts without reading them. Mitigations: show sources, sample-audit sent messages, and track the edit rate by agent. A near-zero edit rate on generated-path drafts is a warning sign.
  • Stale content. Saved replies and help articles drift from policy. Support ops needs a named owner and a review cadence.
  • Language. We don't know the share of non-English tickets. Confirm before launch. If it is significant, scope it out or handle it separately.
  • Legal definition. Does "refund commitment" include credits, prorations, or fee waivers? Legal to define by the end of shadow mode, since it determines G1–G3.
  • Hosting. Confirm data-residency and vendor constraints for ticket content.
  • Threshold ownership. Decide who can change confidence thresholds in production (proposed: support-ops lead, with ML review).

10. Ownership

  • Engineering: integration with the ticketing system, sending-service enforcement (R10), logging, feature flags, latency.
  • ML: data preparation, routing model, drafting pipeline, refund detector, evaluation and monitoring.
  • Support ops: labeling, saved-reply and style-guide upkeep, pilot agents, weekly quality review, and thresholds.
  • Legal: the refund definition, the detector test set sign-off, and audit retention.

Check by check

Got wrong · 1
  • Proposes tests that could failNot every proposed gate has a numeric threshold, read-out window, and specified action; shadow and pilot go/no-go conditions are not fully quantified.
Got right · 9
  • Uses the supplied evidence correctlyAll current-situation facts cited are from the supplied context or derived arithmetically, and causal statements are labelled as uncertain.
  • Addresses the actual decisionThe output commits to a clear PRD design with routing, drafting, rollout and rollback conditions for the named leads.
  • Respects explicit constraintsIt respects the legal no-auto-refund-commitment rule, requires agent approval, targets the named readers, and is within the word limit.
  • Identifies material uncertaintyIt explicitly lists open questions such as wait-vs-handle split, non-English share, legal definition of refund commitment, and hosting constraints, with owners and resolution paths.
  • Avoids unsupported claimsHypotheses such as likely causes of delay are labelled as uncertain, and forecasts are tied to later measurement rather than asserted as fact.
  • Produces the required deliverableThe PRD is complete, practical, reader-appropriate and actionable for engineering, ML and support-ops leads.
  • Specifies AI failure behaviourIt specifies low-confidence routing to General with triage, fallback clarifying drafts when content is insufficient, and draft-timeout behavior.
  • Defines evals and a launch barIt defines an evaluation set, blind review rubric, routing and draft metrics, refund detector recall, and launch bars.
  • Success settles trade-offsIt names target metrics with time frames and gives explicit trade-off rules such as no auto-send at any confidence and low-confidence routing to General.

Grades and run details

Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly9%
  • passAddresses the actual decision77%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims57%
  • passProduces the required deliverable73%
  • passSpecifies AI failure behaviour97%
  • passDefines evals and a launch bar99%
  • partialSuccess settles trade-offs63%
  • partialProposes tests that could fail65%
Run
Run
#1
API response time
46 s
Submitted
29 Sept 2026

Needs more than a quick edit8 of 10 checks passed

Unsupported claims about current workflow and data content would require re-checking evidence, and the data section's assumptions need validation before engineering can act.

Output

PRD: AI Draft & Route for Support Tickets

Sep 24, 2026 · @Martin Slaney

1. Summary and problem

Ledgerly's support team (42 agents, four queues: Billing, Technical, Account access, General) handles \~9,000 tickets a week. Average first response time (FRT) is 7 hours against a target of under 2.

Most of that delay is not writing time. It is triage and lookup1: \~60% of tickets are answered with one of \~40 saved replies, and 38% are billing. Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details.

This feature does two things when a ticket arrives:

  1. Routes it to the correct queue, with a confidence score.
  2. Drafts a first response, grounded in the saved-reply library, the customer's account data and similar past resolutions, for an agent to review.

Nothing is sent without an agent approving it. On top of that, any draft containing a refund commitment is hard-gated per Legal (§4.3).

2. Goals, non-goals and success metrics

Goals

  • Bring average FRT under 2 hours without adding headcount.
  • Cut misrouted tickets and the reassignment delay they cause.
  • Reduce agent handling time on saved-reply tickets, so agents spend more time on the \~40% that need real investigation.

Non-goals (v1)

  • Auto-sending any response, of any type.
  • Drafting follow-ups beyond the first response.
  • Taking actions in Ledgerly (issuing refunds, resetting passwords, changing plans). The model drafts text; agents act.
  • Customer-facing chatbot or deflection.

Success metrics (measured per queue, against a 4-week pre-launch baseline)

MetricTargetGuardrail
Average FRT< 2hP90 FRT must not rise
Routing accuracy (final queue = predicted queue)≥ 92% overall, ≥ 90% per queueAccount access recall ≥ 95%
Draft acceptance (sent with light or no edits)≥ 50% of drafted tickets—
Agent handling time on drafted tickets−30%—
CSAT on drafted ticketsNo drop vs baseline (±1pt)Reopen rate not up >1pt
Refund commitments sent without refund-approval step0Hard requirement

"Light edit" is defined as a normalised edit distance below 0.2 between draft and sent text. ML owns the metric definition; Support Ops signs it off before pilot.

3. Users and core workflow

Users: support agents (reviewers of every draft); queue leads (monitor routing, handle overrides); Support Ops (owns saved replies, policies and the refund-approval rota).

Flow for a new ticket

  1. Ticket created (email or in-app form).
  2. Router predicts queue + confidence within 30s. High confidence → assigned to that queue. Low confidence → Triage view for a lead to assign in one click.
  3. Drafter produces a first response and attaches it to the ticket as an internal draft, with: the saved reply(s) it drew on, account facts it inserted, and any flags (refund, low confidence, missing data).
  4. Agent opens the ticket, reviews the draft, and chooses Send, Edit & send, or Discard (with a reason code).
  5. If the draft is refund-flagged, Send is replaced by Request refund approval (§4.3).
  6. Final queue, sent text and action are logged as training and evaluation signal.

4. Functional requirements

4.1 Routing

  • R1. Classify every new ticket into Billing, Technical, Account access or General, with a calibrated confidence score.
  • R2. Auto-assign when confidence ≥ threshold (set per queue from eval, starting target: ≥ 95% precision at that threshold). Below threshold → Triage view.
  • R3. Account access is the costliest miss (locked-out customers).3 Tune for recall on this class; a ticket with any access signal and ambiguous classification goes to Account access, not General.
  • R4. Agents can reassign in one click; every reassignment is logged with the original prediction.
  • R5. Support Ops can switch routing to suggest-only per queue without a deploy.

4.2 Drafting

  • D1. Generate a draft for every routed ticket in English within 60s of creation. Other languages: no draft in v1, flag only.
  • D2. Retrieval first: identify the best-matching saved reply (or "none"). When one matches, the draft is that reply personalised with ticket and account context, not free text. When none matches, draft from similar resolved tickets and help-centre articles, and label it Free-form draft.
  • D3. Account facts (plan, billing dates, invoice amounts, last payment status) come only from read-only lookups against Ledgerly's billing/account APIs, never from model memory. Any fact the model could not verify is left as a visible `[placeholder]` that blocks sending until filled.
  • D4. Show sources inline: saved-reply ID, linked tickets/articles, API fields used.
  • D5. Never promise timelines, credits, discounts, policy exceptions or refunds unless the saved reply itself contains them2 (refunds additionally gated, §4.3).
  • D6. Skip drafting (flag only) for: legal threats, suspected fraud, data deletion/GDPR requests, security incidents, and abusive or distressed customers. Keyword + classifier; list owned by Support Ops.

4.3 Refund guardrail (Legal requirement)

Legal requires no automated sending of refund commitments. Because every draft already needs agent approval, we go further so a refund promise cannot slip through in a routine approve click:

  • G1. A dedicated refund-commitment detector runs on every draft and on the final edited text at send time (agents may add refund language themselves). Tuned for recall ≥ 99% on a Legal-reviewed test set; false positives are acceptable.
  • G2. If triggered, the ticket cannot be sent via one-click approval. The agent must confirm the refund is authorised under current policy (Support Ops defines who can authorise which amounts) via a separate confirmation step that is logged.
  • G3. The drafter must never generate a refund commitment from free-form reasoning; refund language may only come from approved refund saved replies.
  • G4. No bulk-approve action exists anywhere in the product.
  • G5. Legal reviews the detector test set and the confirmation UX before pilot, and receives a monthly log of refund-flagged sends.

Open for Legal: does "automated sending" cover a one-click approve by an agent? This PRD assumes one-click approval is not automated but adds G2 as defence in depth. Confirm before build.

4.4 Agent experience

  • A1. Draft appears pre-filled in the reply box, visibly marked as AI-drafted until edited.
  • A2. Discard requires a reason: wrong answer, wrong tone, missing info, wrong queue, should not be drafted.
  • A3. Agents never lose their normal tools; saved replies remain available manually.
  • A4. Customers are not told a draft was AI-assisted (agent authors the sent message). Support Ops to confirm this against Ledgerly's AI disclosure policy.

5. ML approach, data and evaluation

Data. Two years of resolved tickets (\~900k) with final queue, agent replies and saved-reply usage.

  • Label routing from the final queue, not the initial one; tickets that were reassigned are the most valuable examples.
  • Map historic replies to saved-reply IDs where possible (exact/near-match), giving a supervised "which reply fits" dataset.
  • Down-weight or exclude replies older than any policy or pricing change; Support Ops supplies the change dates. Stale answers are the main quality risk.
  • Strip PII before any use outside the production data boundary; confirm data-processing terms with Legal/DPO for customer ticket content.

Approach. Start with the simplest thing that hits the targets: an LLM classifier (or fine-tuned small model if cost/latency requires) for routing; retrieval over saved replies + resolved tickets + help centre, then LLM personalisation for drafting; a separate refund detector. Model choice is ML's call, constrained by the latency targets in §3 and D1 and per-ticket cost.

Offline evaluation (gate to pilot)

  • Held-out set: most recent 3 months, time-split (no leakage).
  • Routing: per-queue precision/recall and calibration; auto-assign coverage at the chosen threshold.
  • Drafting: saved-reply selection top-1 accuracy; 500-ticket human-graded sample scored by senior agents (correct, complete, on-policy, tone) with a ≥ 80% "send with light or no edit" bar.
  • Refund detector: recall ≥ 99% on Legal-reviewed set.
  • Factual accuracy: 0 invented account facts in the graded sample.

Online monitoring. Dashboards per queue for acceptance, edit distance, discard reasons, reassignments, FRT, CSAT and refund flags. Weekly review of a random sample of sent drafted replies by Support Ops QA.

6. Rollout, risks and open questions

Rollout

PhaseScopeExit criteria
0. Shadow (2 wks)Route + draft on all tickets, invisible to agentsOffline bars met on live traffic
1. Pilot (3 wks)Billing queue, 8 agents; routing suggest-onlyAcceptance ≥ 40%, no CSAT drop, 0 refund-guardrail misses
2. ExpandAll queues drafting; routing auto-assign above thresholdMetrics in §2 hold for 2 consecutive weeks
3. Steady stateMonthly threshold review, quarterly saved-reply refresh—

Billing goes first: highest volume, most saved-reply coverage, and it exercises the refund guardrail early under close watch.

Risks

  • Rubber-stamping. Agents approve drafts without reading them. Mitigate: placeholders block send, refund gate, QA sampling, track time-on-draft before send.
  • Stale policy in drafts. Mitigate: saved replies as primary source, policy-change exclusion in training data, owner for saved-reply freshness.
  • Misrouting Account access. Mitigate: recall-weighted threshold (R3), Triage fallback.
  • Agent trust. Early bad drafts kill adoption. Mitigate: shadow phase, drafts only above a quality-confidence score; otherwise show the suggested saved reply only.

Open questions

  1. Legal: does agent one-click approval count as "automated sending" (§4.3)?
  2. Support Ops: who can authorise refunds, and at what amounts?
  3. Eng: which helpdesk platform APIs support pre-filled internal drafts and blocking send?
  4. DPO: approval to use historic ticket content with a model provider.
  5. Support Ops: AI-assistance disclosure to customers (A4).

Owners: Eng lead (integration, guardrail enforcement, UI), ML lead (models, eval, monitoring), Support Ops lead (saved replies, policy, pilot, QA).

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisSubstantial rework

    Source: VolumesAverage first response is 7 hours against a target of under 2.

    What we’d changePresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: RiskLegal requires no automated sending of refund commitments.

    What we’d changeState the refund rule once and consistently: generated drafts never contain refund commitments, and any commitment an agent adds goes through the separate confirmation step.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changePresent it as an assumption to check with Support Ops. The brief doesn't rank the cost of misrouting by queue, and the routing design leans on it.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyThe PRD states that the delay cause is triage and lookup and that tickets wait unassigned, and assumes saved-reply usage is in the historical data, none of which is evidence from the brief.
  • Avoids unsupported claimsPresents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
Got right · 8
  • Addresses the actual decisionCommits to building the feature with phased rollout and exit criteria that would halt expansion if not met, framed for engineering/ML/support-ops leads.
  • Respects explicit constraintsRespects all constraints: addresses the intended readers, within word count, mandates agent approval, and enforces no automated refund commitments via guardrails.
  • Identifies material uncertaintyIdentifies open questions about legal definition, refund authorizations, platform APIs, data use, and AI disclosure, with owners and action to resolve.
  • Produces the required deliverableProvides a complete PRD for the required audience, within the 1,000–1,500 word range, that they could act on.
  • Specifies AI failure behaviourDefines low-confidence routing to a Triage view, placeholders blocking send, draft suppression for sensitive cases, and discard reasons.
  • Defines evals and a launch barSpecifies offline evaluation with human-graded sample and 80% send-with-light-edits bar, routing precision/recall thresholds, and a refund-detector recall ≥99%.
  • Success settles trade-offsSets success metrics with targets (FRT <2h, routing accuracy ≥92%) and explicit trade-off rules like prioritizing account access recall at the expense of other queues.
  • Proposes tests that could failEach pilot phase has numeric exit criteria (e.g., acceptance ≥40%, 0 refund misses), measurement windows (3 weeks), and triggers that stop rollout if not met.

Claims the judge couldn’t find in the brief

  • Most of that delay is not writing time. It is triage and lookup
  • Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details
  • Two years of resolved tickets (~900k) with final queue, agent replies and saved-reply usage are available
  • Account access is the costliest miss (locked-out customers)
  • Billing queue has the most saved-reply coverage

Grades and run details

Decision model 75 · LLM judge 8 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly46%
  • passAddresses the actual decision69%
  • passRespects explicit constraints35%
  • passIdentifies material uncertainty87%
  • failAvoids unsupported claims71%
  • passProduces the required deliverable69%
  • passSpecifies AI failure behaviour97%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs80%
  • partialProposes tests that could fail71%
Artefacts
Run
Run
#1
Time to output
2.0 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD