Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.9 · synthetic · AI product, support

Needs more than a quick edit7 of 10 checks passed

It is largely usable but invents current systems not in the brief, so a PM would need to verify the evidence before acting. It also lacks an explicit trade-off rule and complete test windows/actions, requiring more than a quick edit.

Output

PRD: AI-Assisted First Response and Ticket Routing

Product: Ledgerly Support Assist Audience: Engineering, ML and Support Operations leads Status: Proposed MVP; thresholds below are launch targets, not measured results Accountable owner: Support Operations lead, with Engineering and ML leads owning delivery and model quality

1. Problem and objective

Ledgerly’s 42 support agents handle 9,000 tickets per week across Billing, Technical, Account access and General. Billing represents 38% of volume—approximately 3,420 tickets weekly. Average first-response time is seven hours, against a target of under two hours. Approximately 60% of tickets receive one of roughly 40 saved replies.

The feature will classify new tickets into the appropriate queue and prepare a grounded first-response draft for an agent to review. It should reduce queue-selection work and repetitive writing without delegating customer communication or financial commitments to the model.

Every outbound response requires explicit agent approval. Legal prohibits automated sending of refund commitments. The MVP will have no automated-send path and will not generate refund promises; agents must author any commitment using the existing authorized refund process.

2. Goals and success measures

Primary goals:

  • Reduce average time to the first substantive human-approved response to under two hours during a staffed pilot.
  • Reduce median active agent time spent preparing first responses by at least 30%.
  • Route tickets accurately while avoiding hidden misroutes.
  • Preserve response correctness, account security and customer satisfaction.

Measure first-response time from ticket creation to the first substantive response sent by an agent. Automated receipts do not count. Report calendar-hour and staffed-hour results separately, plus median and p90, to prevent averages hiding long waits.

Additional proposed launch and pilot thresholds:

MeasureTarget
Queue-classification accuracy≥95% on the adjudicated holdout
Account access routing recall≥98%
High-confidence automatic-routing precision≥98%
Drafts usable with no substantive correction≥80% in blinded review
Unsupported financial or security commitmentsZero observed in launch evaluation
Draft availability latencyp95 ≤15 seconds after ingestion
Pilot qualityNo material deterioration in QA scores or customer satisfaction

Zero observed errors is a release gate, not proof of zero production risk. ML must report sample sizes and confidence intervals. Support Operations must confirm staffing coverage: drafting improvements alone cannot guarantee the response-time target.

3. Scope

MVP includes:

  • Classification into Billing, Technical, Account access or General.
  • Automatic routing only above validated, queue-specific confidence thresholds.
  • Retrieval of approved saved replies and current support documentation.
  • A first-response draft with internal source references, risk flags and suggested clarifying questions.
  • Agent controls to edit, discard, regenerate, change queue and approve/send.
  • Audit logs, monitoring and immediate disable controls.

Excluded: automatic sending, subsequent-turn assistance, ticket resolution, refunds, account changes, security verification decisions and customer-facing AI chat.

Ticket channels, languages and attachment formats are not specified. Support Operations will inventory them before implementation. MVP supports only validated text channels and languages; unsupported inputs receive normal human triage without a generated draft.

4. User workflow and functional requirements

  1. Ingest: On creation of a new eligible ticket, capture its text, subject and approved metadata. Use only authorized customer/account context available to the assigned support role.
  2. Assess: Predict a queue, calibrated confidence and risk flags such as refund request, account compromise or insufficient information.
  3. Route: Assign high-confidence tickets to the predicted queue. Send uncertain cases to the existing manual-triage destination, provisionally General. Surface uncertainty prominently; do not treat General as a confident classification.
  4. Draft: Retrieve relevant approved material and produce a concise response addressing the request. Prefer adapting an applicable saved reply over open-ended generation.
  5. Review: Show the draft, suggested queue, risk flags and source links in the agent workspace. Clearly label the text “AI draft—not sent.”
  6. Approve/send: The existing send action requires an explicit authenticated agent action. No background job, timeout or model output may trigger sending.
  7. Learn: Record queue corrections, edits, rejection reasons and quality reviews for controlled evaluation and future retraining.

Drafts must not invent transactions, troubleshooting outcomes, refund eligibility or account status. Where evidence is missing, ask for necessary information or acknowledge that an agent must investigate. Avoid requesting passwords, full payment-card details or other unnecessary sensitive data.

Refund-related drafts may acknowledge the request and explain approved next steps, but must not promise an amount, eligibility or processing date. Flag these tickets for Billing review. Agents may add commitments only after authorized verification.

A draft becomes stale if relevant ticket content or account context changes. Disable approval until it is refreshed or explicitly reviewed against the updated context. Manual queue changes take precedence over later model results.

5. Data and ML approach

Two years of resolved tickets and agent replies are available. Historical replies are examples, not authoritative policy: outdated guidance and unauthorized commitments must not be reproduced.

ML and Support Operations will:

  • Define queue labels and rules for mixed-intent tickets. Account-security concerns take precedence over routine billing questions; Support Operations must approve the complete precedence matrix.
  • Audit final queue labels and sample ambiguous cases for expert adjudication.
  • Remove duplicates, signatures and irrelevant quoted history; redact unnecessary personal and payment information.
  • Preserve ticket/thread/customer grouping across splits to reduce leakage.
  • Use a chronological training, validation and held-out test split, with the newest period reserved for testing.
  • Evaluate only information available when the ticket arrived. Later replies and resolutions may provide labels but must never become runtime inputs.

Benchmark saved-reply retrieval and a conventional classifier before introducing more complex models. Select the simplest approach meeting quality, latency and operational requirements.

Generation will use current, versioned approved content. Historical replies may support offline training subject to privacy and quality approval, but must not serve as an unrestricted runtime knowledge source. Conflicting or missing sources trigger a clarification or human-investigation draft rather than a guessed answer.

Customer text and retrieved content are untrusted inputs. Instructions embedded in tickets must not change system rules, authorize actions or bypass review.

6. Engineering design and controls

Implement an asynchronous pipeline behind feature flags:

Ticket event → eligibility check → classifier → routing decision → retrieval → generation → policy validation → draft storage → agent UI.

Persist ticket ID/version, model and prompt versions, retrieved document versions, queue scores, draft state, risk flags and agent actions. Use idempotency keys to prevent duplicate processing and concurrency controls to avoid overwriting agent work.

The AI service must have no credential or permission to send messages, issue refunds or modify accounts. Routing permissions must be limited to approved queues. Outbound messages remain controlled by the existing authenticated support application.

Post-generation checks will block prohibited promises and unsupported sensitive claims. Failed checks suppress the draft and display a reason; they must not merely append a disclaimer.

Apply role-based access, encryption and existing retention rules. Any external model provider requires Security and Legal approval, including contractual restrictions on retention and training use. Log identifiers and operational metadata where possible, not unrestricted ticket bodies.

On timeout, provider outage or validation failure, preserve normal ticket handling. Keep the ticket visible, apply manual triage where necessary and show “Draft unavailable.” Never delay ticket intake while waiting for AI.

7. Evaluation and release gates

Build an adjudicated test set covering all queues, common saved-reply cases, mixed intent, sparse descriptions, refund disputes, account compromise, outdated-policy examples and prompt-injection attempts. Report overall and per-queue performance; oversampled risk cases must also be reported separately from production-weighted results.

Two support reviewers will score drafts for correctness, relevance, completeness, tone and policy compliance, with disagreements adjudicated. Distinguish cosmetic edits from substantive corrections. Track source support and appropriate abstention, not acceptance rate alone.

Required acceptance tests include:

  • A refund-request draft contains no commitment.
  • No model or pipeline component can invoke sending.
  • Unsupported facts cause omission or escalation.
  • Ticket updates invalidate stale drafts.
  • Human rerouting is not overwritten.
  • Provider failure leaves tickets accessible.
  • Duplicate events create neither duplicate drafts nor duplicate sends.
  • Agent permissions also constrain retrieved context.

Production expansion requires meeting quality thresholds and passing all safety-critical tests. Any unauthorized financial promise in evaluation blocks launch pending remediation.

8. Rollout, ownership and operations

Indicative eight-week plan:

  • Weeks 1–2: Support Operations approves taxonomy, saved replies, triage rules and baseline measurements; ML audits data; Engineering completes integration and security design.
  • Weeks 3–4: Build the pipeline, UI, evaluation harness and operational controls.
  • Weeks 5–6: Run offline evaluation and shadow mode without changing queues or showing drafts.
  • Weeks 7–8: Pilot with selected agents, then expand by queue only after gates pass.

Use a randomized holdout where practical, balanced by queue and shift, to compare handling time, response time, corrections and customer outcomes.

Engineering owns availability, access controls and rollback. ML owns evaluation, calibration and drift monitoring. Support Operations owns policy content, QA, agent training and staffing. Legal approves refund wording and provider terms.

Provide separate kill switches for routing and drafting. Suspected unauthorized commitments, data exposure or account-security mishandling trigger immediate suspension of the affected capability and incident review. Review corrections and rejected drafts weekly; retraining requires fresh holdout evaluation, not automatic ingestion of agent edits.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyIt invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
  • Proposes tests that could failSeveral acceptance tests and kill criteria lack numeric thresholds, measurement windows, or the action each outcome triggers.
Mixed · 1
  • Success settles trade-offsIt does not give an explicit trade-off rule such as accepting lower coverage to preserve precision above a stated level.The two graders disagreed on this one.
Got right · 7
  • Addresses the actual decisionIt commits to an MVP with a clear approval flow, release gates, and conditions that would change rollout.
  • Respects explicit constraintsIt respects the named audience, length, human approval requirement, and legal prohibition on automated refund commitments.
  • Identifies material uncertaintyIt names missing channels/languages, unmeasured thresholds, sample-size limitations, and how evaluation and pilot results would resolve them.
  • Avoids unsupported claimsNormative requirements and design rules are framed as proposals, not as established facts from the supplied evidence.
  • Produces the required deliverableIt is a complete PRD for engineering, ML, and support-operations leads that could start work from.
  • Specifies AI failure behaviourIt specifies low-confidence routing to manual triage, abstention or clarification when evidence is missing, stale-draft disabling, and outage fallback.
  • Defines evals and a launch barIt defines an adjudicated test set, quality metrics, thresholds, and launch gates.

Claims the judge couldn’t find in the brief

  • The existing send action requires an explicit authenticated agent action.
  • There is an existing manual-triage destination, provisionally General.
  • There is an existing authorized refund process agents must use.

Grades and run details

Decision model 90 · LLM judge 7 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly7%
  • passAddresses the actual decision90%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty44%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable75%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs26%
  • partialProposes tests that could fail69%
Run
Run
#1
API response time
49 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The PRD is thorough, actionable, and requires at most minor wording edits before engineering, ML, and support-ops leads can start work.

Output

Ledgerly: AI first-response drafts and ticket routing

Status: Proposed MVP requirements. Numerical gates below are proposed targets, not observed performance. Product owns scope; Engineering owns workflow integrity; ML owns model quality; Support Operations owns queue policy, content and review quality.

1. Problem and outcome

Ledgerly’s 42 support agents handle 9,000 tickets weekly across Billing, Technical, Account access and General. Billing represents 38%, approximately 3,420 tickets weekly. Average first response is seven hours against a target below two. Approximately 60% of tickets use one of 40 saved replies.

The feature should reduce triage and composition time without shifting effort into correcting poor drafts or recovering misrouted tickets. Every outgoing draft requires human approval. Legal prohibits automated sending of refund commitments.

Faster drafting alone may not close the five-hour gap. Before piloting, Support Operations must measure arrival-to-assignment, assignment-to-review and review-to-send delays by queue and staffed hours, then identify any coverage or staffing changes needed.

2. Scope and operating boundary

MVP processes newly created English-language, text-based tickets in the existing support workspace. It assigns a queue and prepares a first-response draft. Threads, ticket metadata and approved support content are inputs. Attachments are not interpreted; attachment-dependent or unsupported-language tickets remain available for manual handling.

Excluded: follow-up generation, autonomous sending, refunds, account changes, payment actions and staffing optimisation. Existing spam, security and priority rules run first and cannot be overridden by the model.2

Routing and drafting operate independently. A drafting failure must not prevent routing or agent access. An uncertain route must not prevent a useful draft.

3. Agent workflow

  1. Ticket creation starts asynchronous processing. The ticket is immediately visible and its response clock starts at original receipt.
  2. The system records a route, reason and confidence tier, then prepares a draft where supported.
  3. The agent sees the original message, assigned queue, editable draft, source references and warnings. Sources and warnings are internal only.
  4. The agent can change queue, edit, regenerate, discard or write manually. Regeneration never overwrites unsaved edits without confirmation.
  5. Selecting Approve and send authorises the exact visible text. There is no bulk approval.

Draft states are pending, ready, unavailable, stale and sent. Queue changes preserve receipt time and draft history. New customer messages or changes to relevant ticket context mark drafts stale and require renewed review. Show actionable failure messages and keep the manual composer available.

4. Routing requirements

Support Operations owns this initial taxonomy:

QueuePrimary issue
BillingCharges, subscriptions, invoices, cancellations and refund requests
TechnicalErrors, integrations, imports and malfunctioning features
Account accessLogin, authentication, permissions and suspected account takeover
GeneralProduct guidance and genuinely uncategorised enquiries

For multiple intents, suspected account compromise takes precedence, followed by access-blocking issues; otherwise route by the customer’s main requested resolution. Record secondary intents for the receiving agent.

Automatically assign only when a queue-specific threshold meets the evaluation gate below. Confidence must be calibrated against labelled examples, not taken from a model’s self-reported certainty.

Below threshold, assign to General with a distinct Needs triage status and show the leading suggestions. This is a fallback assignment, not a successful classification. Support Operations assigns a named triage owner each shift and reviews these tickets at least every 30 minutes during staffed hours. Existing out-of-hours escalation remains in force.

Agents can override any route and optionally record a reason. Never automatically reroute after an agent takes ownership. Log overrides for review, not immediate retraining.

5. Draft content and refund controls

Start with retrieval from the approximately 40 saved replies and current, Support Operations-approved help and policy content. Adapt an applicable reply before attempting a novel answer. Each source has an owner, version and review date; withdrawn content becomes unavailable immediately.

Drafts must answer the stated question, request necessary missing details and cite supporting sources internally. They must not invent account facts, troubleshooting outcomes, eligibility, amounts or dates. Account-specific claims require authorised, current account context. Without it, draft a clarification or indicate that an agent must investigate.

For refund requests, MVP drafts may acknowledge the request and explain approved review steps, but must not promise eligibility, payment amounts or payment dates. Agents may manually add commitments only under Ledgerly’s existing refund authority policy. Flag refund-related tickets and require an explicit acknowledgement when the final text contains a detected commitment. This check assists reviewers; it is not the legal enforcement boundary.

The enforcement boundary is the sending service: generation workers have no send credentials. Every AI-assisted send requires an authenticated agent approval bound to ticket ID, recipient and exact message version. Any subsequent edit invalidates approval. Retries cannot send duplicate messages. Existing automation must not consume AI drafts as sendable replies. Therefore, missed refund detection cannot trigger autonomous sending.

6. Data and ML approach

Use the two-year resolved-ticket archive to learn routing patterns and evaluate drafting. Historical agent replies are examples, not authoritative policy. Resolution does not establish correctness or refund permission.

Support Operations must relabel a representative sample against today’s taxonomy, with two reviewers and adjudication for disagreements. Redact credentials, payment details and unnecessary personal information. Preserve tenant boundaries and restrict access to approved training personnel and services.

Split chronologically into training, validation and a locked recent test set. Keep complete conversations and duplicate ticket clusters in one split. At inference and evaluation, expose only information available before the first response; later replies and final queue labels must not leak into inputs.

Begin with a classifier plus retrieval-grounded generation. Fine-tuning is optional and requires measurable improvement over this baseline. No automatic learning from approvals. Provider data use and retention must be approved before production data leaves Ledgerly.

Treat customer text and retrieved content as untrusted data: embedded instructions cannot change policies, access other tenants or invoke privileged actions.

7. Evaluation and launch gates

Build a locked test set of at least 1,000 tickets, stratified across queues, with at least 150 per queue. Include saved-reply matches, ambiguous intents, refunds, outdated policies, missing context, prompt injection and sensitive-data cases. Report volume-weighted results and per-queue results separately.

Required gates:

  • Routing: at least 95% precision among automatically assigned tickets in each queue. Report confidence intervals, recall, coverage and the confusion matrix. Target at least 60% overall automatic-routing coverage; do not lower precision to achieve coverage.
  • Drafts: at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.1 Score factual accuracy, policy compliance, completeness and tone separately.
  • Critical errors: zero observed cross-tenant disclosures, unsupported refund commitments in generated drafts or account-security instructions that bypass verification. Any occurrence blocks release pending correction and regression testing.
  • Workflow: all permission, stale-approval, recipient-change, duplicate-event and retry tests pass. Attempts to send without valid approval must fail.

Use blinded human review and adjudicate disagreements. Automated grading can assist sampling but cannot determine safety gates alone. Zero observed errors is a test result, not proof of zero production risk.

8. Pilot, measurement and operations

Run one week in shadow mode across all queues, followed by a two-week pilot with agents from every queue. Randomise eligible tickets within queue and shift between assisted and existing workflows; account for shared-agent effects when interpreting results.

Primary pilot gate: at least 20% lower mean receipt-to-first-human-response time versus control. Track progress towards the under-two-hour target across all incoming tickets, including unsupported and fallback cases. Also report median, p90, percentage answered within two hours, review time, queue transfers, correction severity, backlog age and customer satisfaction. Draft creation is not a first response.

Expand only if offline gates hold, human handling time improves and transfer rates and customer satisfaction show no material deterioration. Define tolerances and adequate sample sizes before the pilot; extend measurement when inconclusive.

Engineering provides independent routing and drafting kill switches, audit logs and alerts. Target p95 draft readiness within 60 seconds; after timeout, mark unavailable and retain manual handling. Retries must not overwrite agent work.

Support Operations samples 50 assisted tickets weekly, oversampling refunds and overrides. ML monitors quality and drift by queue and content version. Any unauthorised send or cross-tenant disclosure triggers immediate suspension of the affected feature and incident review.

Before pilot launch, Engineering verifies send-path enforcement and integration contracts; ML publishes evaluation results; Support Operations approves taxonomy, sources and shift ownership; Legal confirms the refund workflow.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.

  2. 2Invented evidenceVerify or remove the claimQuick edit

    What we’d changeThe brief doesn't mention existing spam, security or priority rules, or out-of-hours escalation. Confirm they exist before the design relies on them.

Check by check

Mixed · 1
  • Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the supplied context or arithmetic, with no invention.The two graders disagreed on this one.
Got right · 9
  • Addresses the actual decisionThe PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
  • Respects explicit constraintsThe no-automated-refunds requirement is enforced via draft rules, approval checks, and a send-path enforcement boundary; length and reader constraints are met.
  • Identifies material uncertaintyUnknowns like actual delay components, zero-error test limits, and pilot inconclusiveness are named, with resolution steps specified.
  • Avoids unsupported claimsHypotheses (e.g., 'faster drafting alone may not close the gap') are clearly flagged as possibilities, not fact.
  • Produces the required deliverableA complete PRD with scope, workflow, ML approach, evals, and pilot plan is delivered for the target leads, within ~1,200 words.
  • Specifies AI failure behaviourFallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
  • Defines evals and a launch barA 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
  • Success settles trade-offsSuccess metric (≥20% reduction in response time) with a pilot time frame, and an explicit precision-over-coverage trade-off rule are given.
  • Proposes tests that could failAll evaluation and pilot gates have numeric thresholds, a two-week measurement window, and defined actions (block release, suspend feature).

Grades and run details

Decision model 85 · LLM judge 11 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly2%
  • passAddresses the actual decision85%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty83%
  • passAvoids unsupported claims75%
  • passProduces the required deliverable75%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs89%
  • partialProposes tests that could fail73%
Run
Run
#1
Time to output
1.8 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD