Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 42% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The output commits to a clear product decision (generate a reviewable draft, host reviews and publishes, no auto‑emails without review) and explains when it would change (if attendance instrumentation is absent, assignments cannot launch).
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  2. Specifies AI failure behaviour98% pass
    The PRD specifies what happens when the AI is uncertain (block assignment, flag for review, never fall back to calendar), and defines failure states, retries, and transactional guarantees.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items
  3. Produces the required deliverable94% pass
    The output is a complete PRD for engineering and design leads, within the 1,000–1,500 word range, with enough detail to start implementation without major gaps.
    GPT-6 Astra · ChatGPT · Meeting summaries with action items

Where it slips

  1. Uses the supplied evidence correctly29% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail38% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims46% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.9 · synthetic · AI product, support

Usable with a quick edit9 of 10 checks passed

The PRD is usable as-is with only small edits to quantify the pilot go/no-go thresholds; the evidence and reasoning are sound and would not need redoing.

Output

PRD: AI-Drafted First Responses and Ticket Routing

Status: Draft for build planning Audience: Engineering, ML, and Support Operations leads Product: Ledgerly Support

1. Problem

Ledgerly receives about 9,000 support tickets a week, handled by 42 agents across four queues: Billing, Technical, Account access, and General. Billing is 38% of volume (about 3,400 tickets a week). Average first response is 7 hours against a target of under 2.

Two things likely drive the delay, and we don't yet know their split:

  • Wait time. Tickets sit in a queue, or in the wrong queue, before anyone opens them.
  • Handle time. Agents spend time writing or assembling a first reply. About 60% of tickets are already answered with one of roughly 40 saved replies, so much of this work is repetitive.

This feature has two parts. It routes each new ticket to the right queue, and it drafts a first response for the agent to review, edit, and send. Nothing is ever sent without an agent's approval.

2. Goals and non-goals

Goals 1. Bring mean first response time under 2 hours within 90 days of full rollout. 2. Route at least 92% of tickets to the correct queue on first assignment. 3. Have agents send at least 50% of drafts with no or light edits. 4. Ensure no draft or sent message contains an unreviewed refund commitment (see §5).

Non-goals - Auto-sending any message, in any queue, at any confidence level. - Resolving tickets, or taking actions such as issuing refunds, changing plans, or resetting access. - Replacing follow-up replies after the first response (later phase). - Changing queue structure or staffing.

3. Users

  • Support agents (42): review, edit, approve, or discard drafts, and correct routing.
  • Support-ops leads: own saved replies, policy content, thresholds, and quality review.
  • Legal/compliance: consulted on the refund guardrail and audit logs.

4. Functional requirements

4.1 Routing

  • R1. On ticket creation, classify into one of the four queues and assign it automatically.
  • R2. Store a confidence score with each assignment. Below a threshold set by support ops (initially tuned so about 10% of tickets fall below it), assign to General with a "needs triage" tag and show the top two suggested queues.
  • R3. Agents can re-route in one click and give an optional reason. Every re-route is logged as a labeled correction.
  • R4. Also emit tags used by drafting: a topic label (for example, invoice question, failed payment, login lockout) and a refund-related flag.

4.2 Draft generation

  • R5. A draft must be ready in the agent's ticket view within 60 seconds of ticket creation, and never block the ticket from appearing. If drafting fails or times out, the ticket appears as it does today.
  • R6. Use two drafting paths:
  • Saved-reply path. If a ticket matches one of the ~40 saved replies above a similarity threshold, use that reply with variables filled in (name, plan, invoice number). The wording stays as approved by support ops.
  • Generated path. Otherwise, generate a draft grounded in retrieved help-center articles and similar past resolved tickets. Every factual claim about policy, pricing, or product behavior must come from retrieved content. If the content is insufficient, the draft asks the customer a clarifying question or says an agent is investigating, rather than guessing.
  • R7. Each draft shows which path produced it and its sources (saved reply ID or article links) so agents can verify quickly.
  • R8. Match the tone and format of the best-rated historical agent replies. Support ops maintains the style guide.

4.3 Agent workflow

  • R9. Drafts appear pre-filled in the reply composer with three actions: Send (after optional edits), Discard, and Regenerate. Discarding takes an optional reason from a short list.
  • R10. The system has no auto-send code path. Sending requires an authenticated agent action, and this is enforced in the sending service, not just the UI.
  • R11. Log for each ticket: draft text, final sent text, edit distance, agent ID, time from ticket open to send, and discard or re-route reasons.

4.4 Account-access safeguards

  • R12. Drafts for Account access tickets must never confirm whether an account exists, reveal account details, or state that access was restored or changed. They may give standard verification instructions from approved content only.

5. Refund guardrail (legal requirement)

Legal requires no automated sending of refund commitments. Because agents approve every draft, the primary control is R10. We add layered controls so the model never puts a commitment in front of an agent as if it were approved policy.

  • G1. Prompt and template rules. Drafts may acknowledge a refund request and say it is being reviewed. They must not promise, imply, or estimate a refund, credit, waiver, or timeline, and no saved reply used by this feature may contain one.
  • G2. Output classifier. Every draft passes a commitment detector (rules plus model) before display. If it fires, replace the offending sentence with a neutral placeholder, such as "[Agent: refund decision needed]", and tag the ticket.
  • G3. Send-time check. The detector also runs on the final edited text. If an agent's own text contains a refund commitment, Send requires an explicit confirmation ("This message commits to a refund"), which is logged. Agents may make commitments in their own words. The system may not.
  • G4. Audit. Retain drafts and final messages for the period Legal specifies. Support ops reviews a weekly sample of 100 refund-related tickets.
  • G5. Release bar. The detector must reach at least 98% recall on a labeled set of refund-commitment phrasings, including implicit ones like "we'll take care of that charge", before Billing goes live. Legal signs off on the test set.

6. Data and evaluation

Available: two years of resolved tickets with agent replies.

Data preparation (ML lead owns) - Use the queue where a ticket was resolved as the routing label, not where it first landed, since misroutes are the problem we're fixing. Also keep the initial queue to measure the historical misroute rate. - Remove PII before training or indexing. Get security review of the vendor and hosting setup if any external model is used. - Filter out replies that are outdated (old pricing, retired features, superseded policies). Support ops flags policy change dates so the pipeline can exclude earlier replies. - Mark replies that used a saved reply. This gives the saved-reply matching set and a baseline for the other 40%. - Have support ops label about 1,500 recent tickets for refund-commitment presence, routing, and draft quality. This set doubles as the gold evaluation set.

Evaluation - Split by time, not randomly: train on the earliest ~21 months, test on the most recent 3. Random splits will overstate performance because of seasonality and policy drift. - Routing: report accuracy and per-queue precision and recall, plus confusion between Billing and General, and Account access and Technical, which are likely weak spots. - Drafts: blind human review by senior agents on a 5-point rubric (accuracy, policy correctness, tone, completeness), with automatic policy-violation checks. Compare against the original agent reply. - Slice results by queue, topic, and saved-reply versus generated path.

7. Success metrics

MetricBaselineTarget
Mean first response time7h<2h at 90 days post-rollout (<4h at pilot)
Routing accuracy (first assignment)Measure in shadow≥92%
Drafts sent unedited or lightly editedn/a≥50% overall; ≥75% on saved-reply path
Draft discard raten/a<15%
Policy-incorrect drafts in weekly reviewn/a<2%
Refund commitments sent without confirmation flown/a0
Customer satisfaction on first-response ticketsCurrent CSATNo decline

First task: instrument the current split of first-response time into queue wait and agent handle time. Drafting mainly cuts handle time. If wait time dominates, we also need queue-ordering changes, such as surfacing the oldest tickets first, or the 2-hour target will be missed even with high draft quality. Throughput is about 43 tickets per agent per day, so handle time matters, but it is not necessarily the bottleneck.

8. Rollout

  1. Shadow mode (2–3 weeks). Run routing and drafting on live tickets without showing anything to agents. Compare with actual outcomes, set thresholds, and validate the refund detector.
  2. Pilot (3 weeks). Enable routing for all tickets. Show drafts to about 12 agents across General, Technical, and Account access.
  3. Expand. Enable drafts for all agents in those three queues once the pilot meets the quality targets.
  4. Billing. Enable only after G5 is met and Legal has signed off. Start with 6–8 agents, then expand.

Any queue can be switched off by a feature flag within minutes. Rollback triggers: any refund-commitment incident, a policy-incorrect rate above 5% in weekly review, or a CSAT drop of more than 3 points.

9. Risks and open questions

  • Automation bias. Agents may approve drafts without reading them. Mitigations: show sources, sample-audit sent messages, and track the edit rate by agent. A near-zero edit rate on generated-path drafts is a warning sign.
  • Stale content. Saved replies and help articles drift from policy. Support ops needs a named owner and a review cadence.
  • Language. We don't know the share of non-English tickets. Confirm before launch. If it is significant, scope it out or handle it separately.
  • Legal definition. Does "refund commitment" include credits, prorations, or fee waivers? Legal to define by the end of shadow mode, since it determines G1–G3.
  • Hosting. Confirm data-residency and vendor constraints for ticket content.
  • Threshold ownership. Decide who can change confidence thresholds in production (proposed: support-ops lead, with ML review).

10. Ownership

  • Engineering: integration with the ticketing system, sending-service enforcement (R10), logging, feature flags, latency.
  • ML: data preparation, routing model, drafting pipeline, refund detector, evaluation and monitoring.
  • Support ops: labeling, saved-reply and style-guide upkeep, pilot agents, weekly quality review, and thresholds.
  • Legal: the refund definition, the detector test set sign-off, and audit retention.

Check by check

Got wrong · 1
  • Proposes tests that could failNot every proposed gate has a numeric threshold, read-out window, and specified action; shadow and pilot go/no-go conditions are not fully quantified.
Got right · 9
  • Uses the supplied evidence correctlyAll current-situation facts cited are from the supplied context or derived arithmetically, and causal statements are labelled as uncertain.
  • Addresses the actual decisionThe output commits to a clear PRD design with routing, drafting, rollout and rollback conditions for the named leads.
  • Respects explicit constraintsIt respects the legal no-auto-refund-commitment rule, requires agent approval, targets the named readers, and is within the word limit.
  • Identifies material uncertaintyIt explicitly lists open questions such as wait-vs-handle split, non-English share, legal definition of refund commitment, and hosting constraints, with owners and resolution paths.
  • Avoids unsupported claimsHypotheses such as likely causes of delay are labelled as uncertain, and forecasts are tied to later measurement rather than asserted as fact.
  • Produces the required deliverableThe PRD is complete, practical, reader-appropriate and actionable for engineering, ML and support-ops leads.
  • Specifies AI failure behaviourIt specifies low-confidence routing to General with triage, fallback clarifying drafts when content is insufficient, and draft-timeout behavior.
  • Defines evals and a launch barIt defines an evaluation set, blind review rubric, routing and draft metrics, refund detector recall, and launch bars.
  • Success settles trade-offsIt names target metrics with time frames and gives explicit trade-off rules such as no auto-send at any confidence and low-confidence routing to General.

Grades and run details

Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly9%
  • passAddresses the actual decision77%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims57%
  • passProduces the required deliverable73%
  • passSpecifies AI failure behaviour97%
  • passDefines evals and a launch bar99%
  • partialSuccess settles trade-offs63%
  • partialProposes tests that could fail65%
Run
Run
#1
API response time
46 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The PRD is thorough, actionable, and requires at most minor wording edits before engineering, ML, and support-ops leads can start work.

Output

Ledgerly: AI first-response drafts and ticket routing

Status: Proposed MVP requirements. Numerical gates below are proposed targets, not observed performance. Product owns scope; Engineering owns workflow integrity; ML owns model quality; Support Operations owns queue policy, content and review quality.

1. Problem and outcome

Ledgerly’s 42 support agents handle 9,000 tickets weekly across Billing, Technical, Account access and General. Billing represents 38%, approximately 3,420 tickets weekly. Average first response is seven hours against a target below two. Approximately 60% of tickets use one of 40 saved replies.

The feature should reduce triage and composition time without shifting effort into correcting poor drafts or recovering misrouted tickets. Every outgoing draft requires human approval. Legal prohibits automated sending of refund commitments.

Faster drafting alone may not close the five-hour gap. Before piloting, Support Operations must measure arrival-to-assignment, assignment-to-review and review-to-send delays by queue and staffed hours, then identify any coverage or staffing changes needed.

2. Scope and operating boundary

MVP processes newly created English-language, text-based tickets in the existing support workspace. It assigns a queue and prepares a first-response draft. Threads, ticket metadata and approved support content are inputs. Attachments are not interpreted; attachment-dependent or unsupported-language tickets remain available for manual handling.

Excluded: follow-up generation, autonomous sending, refunds, account changes, payment actions and staffing optimisation. Existing spam, security and priority rules run first and cannot be overridden by the model.2

Routing and drafting operate independently. A drafting failure must not prevent routing or agent access. An uncertain route must not prevent a useful draft.

3. Agent workflow

  1. Ticket creation starts asynchronous processing. The ticket is immediately visible and its response clock starts at original receipt.
  2. The system records a route, reason and confidence tier, then prepares a draft where supported.
  3. The agent sees the original message, assigned queue, editable draft, source references and warnings. Sources and warnings are internal only.
  4. The agent can change queue, edit, regenerate, discard or write manually. Regeneration never overwrites unsaved edits without confirmation.
  5. Selecting Approve and send authorises the exact visible text. There is no bulk approval.

Draft states are pending, ready, unavailable, stale and sent. Queue changes preserve receipt time and draft history. New customer messages or changes to relevant ticket context mark drafts stale and require renewed review. Show actionable failure messages and keep the manual composer available.

4. Routing requirements

Support Operations owns this initial taxonomy:

QueuePrimary issue
BillingCharges, subscriptions, invoices, cancellations and refund requests
TechnicalErrors, integrations, imports and malfunctioning features
Account accessLogin, authentication, permissions and suspected account takeover
GeneralProduct guidance and genuinely uncategorised enquiries

For multiple intents, suspected account compromise takes precedence, followed by access-blocking issues; otherwise route by the customer’s main requested resolution. Record secondary intents for the receiving agent.

Automatically assign only when a queue-specific threshold meets the evaluation gate below. Confidence must be calibrated against labelled examples, not taken from a model’s self-reported certainty.

Below threshold, assign to General with a distinct Needs triage status and show the leading suggestions. This is a fallback assignment, not a successful classification. Support Operations assigns a named triage owner each shift and reviews these tickets at least every 30 minutes during staffed hours. Existing out-of-hours escalation remains in force.

Agents can override any route and optionally record a reason. Never automatically reroute after an agent takes ownership. Log overrides for review, not immediate retraining.

5. Draft content and refund controls

Start with retrieval from the approximately 40 saved replies and current, Support Operations-approved help and policy content. Adapt an applicable reply before attempting a novel answer. Each source has an owner, version and review date; withdrawn content becomes unavailable immediately.

Drafts must answer the stated question, request necessary missing details and cite supporting sources internally. They must not invent account facts, troubleshooting outcomes, eligibility, amounts or dates. Account-specific claims require authorised, current account context. Without it, draft a clarification or indicate that an agent must investigate.

For refund requests, MVP drafts may acknowledge the request and explain approved review steps, but must not promise eligibility, payment amounts or payment dates. Agents may manually add commitments only under Ledgerly’s existing refund authority policy. Flag refund-related tickets and require an explicit acknowledgement when the final text contains a detected commitment. This check assists reviewers; it is not the legal enforcement boundary.

The enforcement boundary is the sending service: generation workers have no send credentials. Every AI-assisted send requires an authenticated agent approval bound to ticket ID, recipient and exact message version. Any subsequent edit invalidates approval. Retries cannot send duplicate messages. Existing automation must not consume AI drafts as sendable replies. Therefore, missed refund detection cannot trigger autonomous sending.

6. Data and ML approach

Use the two-year resolved-ticket archive to learn routing patterns and evaluate drafting. Historical agent replies are examples, not authoritative policy. Resolution does not establish correctness or refund permission.

Support Operations must relabel a representative sample against today’s taxonomy, with two reviewers and adjudication for disagreements. Redact credentials, payment details and unnecessary personal information. Preserve tenant boundaries and restrict access to approved training personnel and services.

Split chronologically into training, validation and a locked recent test set. Keep complete conversations and duplicate ticket clusters in one split. At inference and evaluation, expose only information available before the first response; later replies and final queue labels must not leak into inputs.

Begin with a classifier plus retrieval-grounded generation. Fine-tuning is optional and requires measurable improvement over this baseline. No automatic learning from approvals. Provider data use and retention must be approved before production data leaves Ledgerly.

Treat customer text and retrieved content as untrusted data: embedded instructions cannot change policies, access other tenants or invoke privileged actions.

7. Evaluation and launch gates

Build a locked test set of at least 1,000 tickets, stratified across queues, with at least 150 per queue. Include saved-reply matches, ambiguous intents, refunds, outdated policies, missing context, prompt injection and sensitive-data cases. Report volume-weighted results and per-queue results separately.

Required gates:

  • Routing: at least 95% precision among automatically assigned tickets in each queue. Report confidence intervals, recall, coverage and the confusion matrix. Target at least 60% overall automatic-routing coverage; do not lower precision to achieve coverage.
  • Drafts: at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.1 Score factual accuracy, policy compliance, completeness and tone separately.
  • Critical errors: zero observed cross-tenant disclosures, unsupported refund commitments in generated drafts or account-security instructions that bypass verification. Any occurrence blocks release pending correction and regression testing.
  • Workflow: all permission, stale-approval, recipient-change, duplicate-event and retry tests pass. Attempts to send without valid approval must fail.

Use blinded human review and adjudicate disagreements. Automated grading can assist sampling but cannot determine safety gates alone. Zero observed errors is a test result, not proof of zero production risk.

8. Pilot, measurement and operations

Run one week in shadow mode across all queues, followed by a two-week pilot with agents from every queue. Randomise eligible tickets within queue and shift between assisted and existing workflows; account for shared-agent effects when interpreting results.

Primary pilot gate: at least 20% lower mean receipt-to-first-human-response time versus control. Track progress towards the under-two-hour target across all incoming tickets, including unsupported and fallback cases. Also report median, p90, percentage answered within two hours, review time, queue transfers, correction severity, backlog age and customer satisfaction. Draft creation is not a first response.

Expand only if offline gates hold, human handling time improves and transfer rates and customer satisfaction show no material deterioration. Define tolerances and adequate sample sizes before the pilot; extend measurement when inconclusive.

Engineering provides independent routing and drafting kill switches, audit logs and alerts. Target p95 draft readiness within 60 seconds; after timeout, mark unavailable and retain manual handling. Retries must not overwrite agent work.

Support Operations samples 50 assisted tickets weekly, oversampling refunds and overrides. ML monitors quality and drift by queue and content version. Any unauthorised send or cross-tenant disclosure triggers immediate suspension of the affected feature and incident review.

Before pilot launch, Engineering verifies send-path enforcement and integration contracts; ML publishes evaluation results; Support Operations approves taxonomy, sources and shift ownership; Legal confirms the refund workflow.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.

  2. 2Invented evidenceVerify or remove the claimQuick edit

    What we’d changeThe brief doesn't mention existing spam, security or priority rules, or out-of-hours escalation. Confirm they exist before the design relies on them.

Check by check

Mixed · 1
  • Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the supplied context or arithmetic, with no invention.The two graders disagreed on this one.
Got right · 9
  • Addresses the actual decisionThe PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
  • Respects explicit constraintsThe no-automated-refunds requirement is enforced via draft rules, approval checks, and a send-path enforcement boundary; length and reader constraints are met.
  • Identifies material uncertaintyUnknowns like actual delay components, zero-error test limits, and pilot inconclusiveness are named, with resolution steps specified.
  • Avoids unsupported claimsHypotheses (e.g., 'faster drafting alone may not close the gap') are clearly flagged as possibilities, not fact.
  • Produces the required deliverableA complete PRD with scope, workflow, ML approach, evals, and pilot plan is delivered for the target leads, within ~1,200 words.
  • Specifies AI failure behaviourFallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
  • Defines evals and a launch barA 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
  • Success settles trade-offsSuccess metric (≥20% reduction in response time) with a pilot time frame, and an explicit precision-over-coverage trade-off rule are given.
  • Proposes tests that could failAll evaluation and pilot gates have numeric thresholds, a two-week measurement window, and defined actions (block release, suspend feature).

Grades and run details

Decision model 85 · LLM judge 11 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly2%
  • passAddresses the actual decision85%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty83%
  • passAvoids unsupported claims75%
  • passProduces the required deliverable75%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs89%
  • partialProposes tests that could fail73%
Run
Run
#1
Time to output
1.8 min
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.7100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.361.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini50.038.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD