Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.9 · synthetic · AI product, support

Usable with a quick edit9 of 10 checks passed

The PRD is complete and actionable for the named leads, requiring only the addition of explicit numeric thresholds for already-defined gates. No evidence or reasoning would need to be redone.

Output

PRD: AI First-Response Drafting and Ticket Routing

Status: Proposed Product: Ledgerly Support Audience: Engineering, ML, and Support Operations leads

1. Summary

Ledgerly receives about 9,000 support tickets per week. Tickets are handled by 42 agents across Billing, Technical, Account access, and General. Average first response time is seven hours, against a target of under two hours. About 60% of tickets receive one of roughly 40 saved replies.

This feature will recommend a queue and prepare a first-response draft when a new ticket arrives. An agent must review and approve every draft before it is sent. The feature is intended to reduce time spent triaging and composing routine responses, while leaving decisions and customer communication under agent control. It must never send a response autonomously or make a refund commitment on Ledgerly’s behalf.

2. Goals and non-goals

Goals - Reduce median time from ticket creation to first response, with an operational goal of under two hours. - Reduce agent effort on routine first responses by using appropriate saved replies and ticket-specific details. - Route tickets to the best-fit queue, while making uncertain recommendations visible and easy to correct. - Preserve agent review and control over every customer-facing response.

Non-goals - Automatically send, resolve, or close tickets. - Decide refund eligibility, promise a refund, or commit to refund timing. - Draft replies to later messages in an existing conversation in the initial release. - Replace queue ownership, escalation policies, or agents’ judgment.

3. Users and workflow

The primary user is a support agent reviewing new tickets in their assigned queue. Support Operations owns queue definitions, saved replies, and handling guidance. ML and Engineering own model quality, serving, integrations, and monitoring.

For each new ticket, the system will: 1. Read the ticket’s permitted content and available support context. 2. Recommend one of the four queues, with a confidence indicator and brief reason. 3. Generate a first-response draft, preferentially based on a relevant approved saved reply where appropriate. 4. Display the recommendation and draft in the agent workflow. 5. Let the agent edit, discard, or approve. Approval sends the response through the existing support system; no response is sent before approval.1 6. Record the final queue, draft disposition, edits, and outcome for evaluation.

Agents may change the queue before or after reviewing the draft. A draft must remain clearly marked as AI-generated until approved.

4. Functional requirements

Routing

  • Classify each ticket as Billing, Technical, Account access, or General.
  • Show the recommended queue, confidence, and a short explanation grounded in ticket content.
  • Allow agents to override the recommendation; the override becomes an evaluation signal, not an automatic training label.
  • Route low-confidence or ambiguous cases to General and flag them for review. Support Operations must set and approve the launch confidence threshold using offline and pilot results.
  • Do not infer urgency or bypass existing escalation rules in the initial release.

Drafting

  • Generate a concise, relevant first response based on the ticket and approved support materials available to the system.
  • Use a saved reply when it fits, adapting only with verified ticket details. Do not fabricate account status, actions taken, policy, or resolution.
  • When information is insufficient, ask a clear clarifying question or provide a safe acknowledgement rather than guessing.
  • Support agents can edit, discard, or approve the draft. No draft may be sent without an explicit agent action.
  • If the request concerns a refund, the draft must not promise, confirm, or imply a refund or refund timing. It should use approved noncommittal wording and leave eligibility and timing to an agent. Refund-related tickets should be visibly flagged for agent attention.

5. Data, model, and system requirements

Ledgerly has two years of resolved tickets and agent replies. ML should assess data coverage and quality before training: queue labels, duplicate or reopened cases, outdated answers, agent-specific wording, and tickets whose resolution depended on account information unavailable at ticket creation. Historical agent replies are examples, not policy; approved saved replies and current support guidance take precedence.

Create a time-based training, validation, and test split to reduce leakage from repeated tickets or changing policies. Evaluate routing and draft quality separately, including by queue, ticket type, and relevant risk category. Do not train on post-response information when evaluating first-response behavior. Remove or protect unnecessary sensitive data, and use only the customer and account context needed for support.

The serving path must retrieve only context the agent is authorized to see. Engineering must confirm integration points with the ticketing system, queue assignment, saved replies, and send/approval controls before implementation. Log model/version, recommendation, confidence, draft, agent edits, approval or discard, final queue, and response timestamps, subject to Ledgerly’s retention and access policies. Restrict access to these logs and avoid storing secrets or unnecessary personal data in model prompts.

6. Quality, safety, and success measures

Establish a pre-launch baseline and report weekly: - Median and 90th-percentile time to first response; proportion answered within two hours. - Queue recommendation accuracy and agent override rate, overall and by queue. - Draft approval rate, edit rate, discard rate, and time from ticket opening to send. - Agent-rated usefulness and sampled quality review results. - Refund-policy violations, fabricated claims, and other serious customer-impacting errors.

Suggested launch gates, to be confirmed by Support Operations and Legal: routing performance must be no worse than the current human baseline on the held-out set; pilot draft samples must meet an agreed quality rubric; and no refund commitment may appear in an approved draft during safety testing. Any refund commitment, unauthorized send, or material fabricated claim triggers immediate disablement of the affected capability and incident review. Do not optimize approval rate alone: agents must be able to reject drafts without penalty.

7. Rollout and operations

  1. Discovery and offline evaluation: Validate data, current queue definitions, approved saved replies, policy ownership, and integration feasibility. Build a labeled test set reviewed by Support Operations, with explicit refund and other high-risk examples.
  2. Shadow mode: Run recommendations without showing or applying them. Compare predicted queues and drafts with agent outcomes; inspect errors and subgroup performance.
  3. Limited pilot: Enable the workflow for a small, representative group of agents and ticket types. Keep agent approval mandatory. Support Operations reviews a daily sample and gathers structured agent feedback.
  4. Expansion: Expand to all 42 agents only after launch gates are met and Legal and Support Operations approve. Roll back by disabling drafting or routing independently if either degrades.

Support Operations owns the queue taxonomy, approved reply content, quality rubric, and agent training. ML owns evaluation, thresholds, drift monitoring, and model updates. Engineering owns access controls, reliable delivery, audit logs, feature flags, and rollback. A named on-call owner must be assigned before pilot launch. Review quality and performance weekly during rollout and after material model, policy, or saved-reply changes. Re-evaluate against a fresh sample regularly; do not promote agent edits to training data without review and policy checks.

8. Open decisions before build

Confirm the ticketing platform and integration surfaces; which customer/account fields are available at ticket creation; the authoritative source for current support policy; Legal’s approved refund language and review process; the minimum acceptable routing and draft-quality thresholds; and the retention period for prompts, drafts, and audit logs. These decisions are launch prerequisites, not reasons to relax agent approval or refund safeguards.

What a PM had to fix

  1. 1Test or gate too weakTighten the testTargeted repair

    What we’d changeAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.

Check by check

Got wrong · 1
  • Proposes tests that could failProposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
Got right · 9
  • Uses the supplied evidence correctlyAll current-state figures used come directly from the supplied context, with no invented facts about the current situation.
  • Addresses the actual decisionThe PRD commits unambiguously to AI-assisted queue recommendation and draft generation with mandatory agent approval, and specifies conditions for rollback or disablement.
  • Respects explicit constraintsIt respects the named audience, length, PRD form, and enforces the no-automated-refund-commitment constraint through mandatory approval, noncommittal refund wording, and flagging.
  • Identifies material uncertaintyIt identifies material unknowns such as data quality, integration points, policy ownership, and confidence thresholds, and says how to resolve them through evaluation and piloting.
  • Avoids unsupported claimsNo causes, forecasts, or interpretations are presented as established fact; uncertain items are framed as suggested, open, or to be confirmed.
  • Produces the required deliverableThe requested PRD is present, complete, written for engineering, ML, and support-operations leads, and within the requested length.
  • Specifies AI failure behaviourThe PRD specifies low-confidence routing to General with flags, clarifying questions when information is insufficient, and immediate disablement for serious errors.
  • Defines evals and a launch barIt defines a pre-launch baseline, metrics, a held-out test set, launch gates, and a zero-refund-commitment safety condition.
  • Success settles trade-offsIt names a success metric with an under-two-hours target and weekly cadence, and gives trade-off rules such as disabling the affected capability on any refund commitment or unauthorized send.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly23%
  • passAddresses the actual decision71%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty90%
  • passAvoids unsupported claims33%
  • passProduces the required deliverable46%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar44%
  • partialSuccess settles trade-offs85%
  • partialProposes tests that could fail82%
Run
Run
#1
API response time
30 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD