Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given3 items: Scenario, Volumes, Risk
ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer doesThe answer key the graders mark against

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score at 40)
  • Allows automated refund commitments
Case

Staff-level · v1.11 · synthetic · AI product, support

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 10 checks passed

A PM would need to remove or verify invented current-state facts and tighten the refund guardrail and length before using it. The reasoning is otherwise strong but not safe to ship as-is.

Output

Product Requirements Document (PRD)

Feature: AssistRoute (Automated Queue Routing & First-Response Drafting) Target Delivery: Q3 Status: Ready for Engineering Review Owners: Engineering Lead, ML Lead, Support Operations Lead

---

1. Problem Statement & Business Opportunity

Ledgerly processes 9,000 support tickets per week with a team of 42 agents across four queues: Billing (38%), Technical, Account Access, and General.

  • Current FRT: 7.1 hours average first response time.
  • Target FRT: < 2.0 hours.
  • Manual Bottlenecks: Every incoming ticket requires manual reading, classification, and initial triage. Furthermore, approximately 60% of all initial inquiries map cleanly to one of ~40 canonical saved replies (macros), yet agents spend minutes manually searching, adapting, and populating them.

Solution Overview

AssistRoute is a two-part machine learning pipeline that acts on incoming tickets: 1. Predictive Routing: Automatically assigns incoming tickets to one of the four queues. 2. First-Response Drafting: Generates an editable, context-aware draft response pre-populated in the agent’s console before ticket open.

Core Invariant: Human-in-the-Loop (HITL). No draft is ever sent directly to a customer. Support agents retain 100% send authority.

---

2. Key Objectives & Metrics

Metric CategoryTargetMeasurement Method
:---:---:---
First Response Time (FRT)$\le 2.0\text{ hours}$ (across all queues)Timestamp delta: `ticket.createdat` to `firstagentmessage.sentat`.
Routing Accuracy$\ge 93\%$ overall accuracyEvaluated against tickets reassigned to a different queue within 24h.
Billing Routing Accuracy$\ge 95\%$ precision/recallSpecific tracking on Billing due to volume (38%).
Draft Utilization Rate$\ge 65\%$ of ticketsProportion of first responses where the agent accepts the draft (as-is or edited).
Draft Edit Distance$\le 30\%$ Levenshtein edit distanceMeasures generation quality on accepted drafts.
Zero Refund Violation0 instancesHard constraint: Zero drafts promising or confirming refunds reach customers without human authorization.

---

3. Scope & Non-Goals

In Scope

  • Asynchronous classification and routing pipeline for all newly created tickets via email and web form.
  • Generation of a single personalized initial draft based on historical resolutions and the 40 standard macros.
  • Guardrail pipeline strictly forbidding autonomous refund/credit commitments.
  • Feedback loop telemetry (tracking agent accepts, edits, discards, and re-routes).

Out of Scope (Phase 1)

  • Autonomous sending (auto-resolution without human click).
  • Chat/live messaging support (limited strictly to asynchronous ticketing channels).
  • Processing multi-turn responses (AssistRoute generates first responses only).
  • Multi-language support (English only).

---

4. System Architecture & Workflow

``` [Customer Submits Ticket] │ ▼ [Event Ingestion: Webhook] ──▶ [PII Masking & Sanitization] │ ┌───────────────────────┴───────────────────────┐ ▼ ▼ [Routing Classifier] [Draft Generation Pipeline] │ │ Confidence $\ge$ Threshold? │ ├── Yes ──▶ Assign Queue 1. Macro Match / Retrieval └── No ──▶ Assign "General" (Flagged) 2. Prompt Compilation (LLM) │ 3. Deterministic Policy Guardrails │ │ └───────────────────────┬───────────────────────┘ ▼ [Agent Console: Pre-populated View] ├── Action A: Accept & Send ├── Action B: Edit & Send ├── Action C: Discard Draft └── Action D: Re-assign Queue ```

End-to-End Latency SLA

From `ticket.created` webhook ingestion to draft persistence in the ticketing database: $\le 5.0$ seconds (P95).

---

5. Functional Requirements

5.1 Routing Engine (ML Service)

  • FR-1.1: The model must classify each ticket into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • FR-1.2: Confidence Scoring:
  • If `confidence >= 0.85`: Automatically set the ticket's `queue_id` attribute.
  • If `confidence < 0.85`: Set `queueid` to `General` and append the tag `needstriage`.
  • FR-1.3: The Routing Engine must evaluate and route the ticket before generating the draft, as draft prompts require queue-specific system contexts.

5.2 Retrieval & Generation Pipeline

  • FR-2.1 (Context Hydration): The generation service must pull:
  • The ticket subject and body.
  • The authenticated user's metadata: Plan tier (`Solo`, `Growth`, `Enterprise`), account age, and active add-on modules.
  • The closest semantic match from Ledgerly's 40 standard macros.
  • FR-2.2 (Draft Generation): The system must generate a friendly, concise first response that adheres to Ledgerly brand voice, incorporates the customer's name, and directly addresses the primary issue using the retrieved macro logic.
  • FR-2.3 (Confidence Gating): If the model's semantic similarity score against approved knowledge bases or historical solutions falls below `0.70`, no draft shall be rendered. The UI must show: "Draft unavailable: Low context confidence."

5.3 Legal & Compliance Guardrail (Refund Protection)

* FR-3.1 (Deterministic Regex & Semantic Filter): Every draft must pass through a two-stage financial commitment filter prior to saving: 1. Lexical Check: Negative keyword blocking: `refund`, `credit`, `reimburse`, `waive fee`, `reverse charge`, `money back`, `comp`. 2. Semantic Entailment Check: Verification that the text does not commit Ledgerly to monetary compensation or liability. * FR-3.2 (Violation Action): If a draft contains a prohibited financial commitment: * The drafting engine must automatically strip the phrase and substitute safe boilerplate: "I have escalated your billing inquiry to our finance specialists for immediate review." * The draft must display a visible amber badge in the Agent UI: "Financial Review Required — No Commitments Permitted."

5.4 Support Agent Console UX

  • FR-4.1 (Draft Presentation): The draft must render inside the native response editor as an uncommitted state (distinct background styling or watermarked).
  • FR-4.2 (Interaction Actions):
  • Tab / Single-Click Acceptance: Agent hits `Tab` or clicks `Use Draft` to copy the generated text into the active editor.
  • Direct Editing: Editing the draft directly must be seamless without extra confirmation modals.
  • Clear / Reject: Single-click action to clear the generated text and open a blank response field.
  • FR-4.3 (Metadata Panel): A sidebar panel must show:
  • Queue classification and confidence score.
  • Which of the 40 macros informed the draft.
  • A one-click queue reassignment dropdown.

``` +-------------------------------------------------------------------------+ | Ticket #10492: "Charged twice for Payroll add-on" | | Sender: Sarah Jenkins (Growth Tier) | Queue: [Billing (96%)] [Re-route] | +-------------------------------------------------------------------------+ | [ AI Drafted Response ] [Macro Source: Billing_Dup]| | | | Hi Sarah, | | | | Thanks for reaching out. I see that your account was debited twice | | for the Payroll add-on on October 1st. | | | | I have escalated this to our payment operations team to review the | | duplicate transaction details. We will update you here as soon as this | | review is complete. | | | | [ Accept Draft (Tab) ] [ Edit Draft ] [ Reject ] | +-------------------------------------------------------------------------+ | [ Agent Response Box ] | +-------------------------------------------------------------------------+ ```

---

6. ML & Data Engineering Specifications

6.1 Data Preparation (2-Year Historical Corpus)

1. Sanitization: Strip all historical PII (tax identifiers, SSNs, credit card numbers, passwords) using Microsoft Presidio or an equivalent NER pipeline before training/indexing. 2. Filtering: * Drop all historical tickets that required more than 4 re-routes (noisy labels). * Filter out tickets closed with negative customer satisfaction (CSAT $\le 2$). * Exclude responses superseded by outdated accounting rules or old pricing tiers (Ops team to define date cutoffs). 3. Macro Ground-Truth Alignment: Map the 40 canonical macros against historical agent responses to serve as gold-standard reference pairs.

6.2 Model Specifications

  • Routing Classifier:
  • Architecture: Fine-tuned lightweight encoder (e.g., `modern-bert-base` or `RoBERTa-base`) or an optimized classification endpoint.
  • Input: Ticket Subject + Body.
  • Output: Softmax distribution over `[Billing, Technical, Account Access, General]`.
  • Latency Target: $< 200\text{ ms}$.
  • Generative Drafting Model:
  • Architecture: Hosted LLM (e.g., Claude 3.5 Sonnet or GPT-4o-mini) via secure enterprise API with zero-data-retention agreements.
  • Prompt Design: System prompt containing Ledgerly tone guidelines, user context JSON, the selected macro instructions, and strict instructions forbidding financial promises.
  • Temperature: `0.1` (low variability, high determinism).

6.3 Telemetry & Event Logging

The frontend and backend must emit the following events to the analytics lakehouse: * `ticketrouted`: `{ ticketid, predictedqueue, confidence, autoassigned: bool }` * `ticketrerouted`: `{ ticketid, oldqueue, newqueue, agentid }` * `draftgenerated`: `{ ticketid, macroid, promptversion, modelid, generationtimems }` * `draftactioned`: * `action`: `ACCEPTEDASIS` | `EDITED` | `REJECTED` * `originaldraft`: string * `finalsenttext`: string * `levenshteindistance`: int * `agentid`: string

---

7. Safety, Compliance, and Security

  1. Refund Protection Hard Filter: As specified in Section 5.3, any draft containing automated refund promises is blocked from the UI by an air-gapped deterministic validation check that runs after LLM generation and cannot be bypassed by prompt engineering.
  2. Zero Training on Customer Data: LLM vendors must have explicit Data Processing Agreements (DPAs) confirming customer ticket content is not used to train foundation models.
  3. Data Storage & SOC 2: Masked inputs and outputs stored for model evaluation must adhere to standard data retention policies (auto-purged after 90 days).

---

8. Rollout Plan & Milestones

``` Phase 0: Offline Validation & Guardrail Auditing (Weeks 1-3) ├── Train routing model on historical data; evaluate against held-out test set. ├── Run 10,000 synthetic adversarial prompts targeting the refund guardrail. └── Success Gate: Routing accuracy > 93%, 0 adversarial refund breaches.

Phase 1: Shadow Mode (Weeks 4-5) ├── Pipeline processes live tickets; outputs are logged to Databricks/Snowflake. ├── No UI visible to agents. └── Evaluate: Model inference latency, real-time routing accuracy vs. manual triage.

Phase 2: Single-Queue Canary Pilot (Weeks 6-7) ├── Deploy exclusively to the Billing Queue (10 designated agents). ├── Daily sync with Support Operations leads to review draft quality. └── Success Gate: Draft utilization > 50%, no severe agent blockers.

Phase 3: General Availability (GA) & Iteration (Weeks 8+) ├── Enable routing and drafting across all 42 agents and 4 queues. ├── Implement weekly retraining cadence for classifier based on override logs. └── Track progress toward < 2.0 hour FRT target. ```

---

9. Risks & Mitigations

RiskImpactLikelihoodMitigation Strategy
:---:---:---:---
Agent over-reliance (rubber-stamping bad drafts)HighMediumRandom 2% QA audits on accepted drafts by Support Operations leads. Prominent UI reminders that agents are solely responsible for sent text.
Hallucinated accounting/tax guidanceCriticalLowRestrict generation strictly to standard macros and product knowledge base; temperature set to 0.1; fallback to "Draft unavailable" when query context is ambiguous.
Model drift due to product updatesMediumHighWeekly regression testing on macro sets. Support Ops manages a macro repository that directly updates retrieval indexes without code deploys.
Routing loop (agent re-routes ticket repeatedly)LowLowHard ceiling: If a ticket is manually re-routed twice, AssistRoute disables auto-routing for that ticket and locks it to manual handling.

Claims the judge couldn’t find in the brief

  • Current average first response time is 7.1 hours.
  • Every incoming ticket requires manual reading, classification, and initial triage.
  • Agents spend minutes manually searching, adapting, and populating saved replies.
  • New tickets arrive via email and web form.
  • Authenticated user metadata includes plan tier (Solo, Growth, Enterprise), account age, and active add-on modules.
  • Historical tickets contain PII such as tax identifiers, SSNs, credit card numbers, and passwords.
  • Historical tickets include CSAT scores and re-route counts.
  • The company uses Databricks/Snowflake and has SOC 2 controls.

Grades and run details

Decision model 65 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly59%
  • passAddresses the actual decision87%
  • partialRespects explicit constraints26%
  • partialIdentifies material uncertainty34%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable55%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar92%
  • partialSuccess settles trade-offs75%
  • partialProposes tests that could fail69%
Run
Run
#1
API response time
33 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 10 checks passed

The PRD is usable as-is with only small edits to quantify the pilot go/no-go thresholds; the evidence and reasoning are sound and would not need redoing.

Output

PRD: AI-Drafted First Responses and Ticket Routing

Status: Draft for build planning Audience: Engineering, ML, and Support Operations leads Product: Ledgerly Support

1. Problem

Ledgerly receives about 9,000 support tickets a week, handled by 42 agents across four queues: Billing, Technical, Account access, and General. Billing is 38% of volume (about 3,400 tickets a week). Average first response is 7 hours against a target of under 2.

Two things likely drive the delay, and we don't yet know their split:

  • Wait time. Tickets sit in a queue, or in the wrong queue, before anyone opens them.
  • Handle time. Agents spend time writing or assembling a first reply. About 60% of tickets are already answered with one of roughly 40 saved replies, so much of this work is repetitive.

This feature has two parts. It routes each new ticket to the right queue, and it drafts a first response for the agent to review, edit, and send. Nothing is ever sent without an agent's approval.

2. Goals and non-goals

Goals 1. Bring mean first response time under 2 hours within 90 days of full rollout. 2. Route at least 92% of tickets to the correct queue on first assignment. 3. Have agents send at least 50% of drafts with no or light edits. 4. Ensure no draft or sent message contains an unreviewed refund commitment (see §5).

Non-goals - Auto-sending any message, in any queue, at any confidence level. - Resolving tickets, or taking actions such as issuing refunds, changing plans, or resetting access. - Replacing follow-up replies after the first response (later phase). - Changing queue structure or staffing.

3. Users

  • Support agents (42): review, edit, approve, or discard drafts, and correct routing.
  • Support-ops leads: own saved replies, policy content, thresholds, and quality review.
  • Legal/compliance: consulted on the refund guardrail and audit logs.

4. Functional requirements

4.1 Routing

  • R1. On ticket creation, classify into one of the four queues and assign it automatically.
  • R2. Store a confidence score with each assignment. Below a threshold set by support ops (initially tuned so about 10% of tickets fall below it), assign to General with a "needs triage" tag and show the top two suggested queues.
  • R3. Agents can re-route in one click and give an optional reason. Every re-route is logged as a labeled correction.
  • R4. Also emit tags used by drafting: a topic label (for example, invoice question, failed payment, login lockout) and a refund-related flag.

4.2 Draft generation

  • R5. A draft must be ready in the agent's ticket view within 60 seconds of ticket creation, and never block the ticket from appearing. If drafting fails or times out, the ticket appears as it does today.
  • R6. Use two drafting paths:
  • Saved-reply path. If a ticket matches one of the ~40 saved replies above a similarity threshold, use that reply with variables filled in (name, plan, invoice number). The wording stays as approved by support ops.
  • Generated path. Otherwise, generate a draft grounded in retrieved help-center articles and similar past resolved tickets. Every factual claim about policy, pricing, or product behavior must come from retrieved content. If the content is insufficient, the draft asks the customer a clarifying question or says an agent is investigating, rather than guessing.
  • R7. Each draft shows which path produced it and its sources (saved reply ID or article links) so agents can verify quickly.
  • R8. Match the tone and format of the best-rated historical agent replies. Support ops maintains the style guide.

4.3 Agent workflow

  • R9. Drafts appear pre-filled in the reply composer with three actions: Send (after optional edits), Discard, and Regenerate. Discarding takes an optional reason from a short list.
  • R10. The system has no auto-send code path. Sending requires an authenticated agent action, and this is enforced in the sending service, not just the UI.
  • R11. Log for each ticket: draft text, final sent text, edit distance, agent ID, time from ticket open to send, and discard or re-route reasons.

4.4 Account-access safeguards

  • R12. Drafts for Account access tickets must never confirm whether an account exists, reveal account details, or state that access was restored or changed. They may give standard verification instructions from approved content only.

5. Refund guardrail (legal requirement)

Legal requires no automated sending of refund commitments. Because agents approve every draft, the primary control is R10. We add layered controls so the model never puts a commitment in front of an agent as if it were approved policy.

  • G1. Prompt and template rules. Drafts may acknowledge a refund request and say it is being reviewed. They must not promise, imply, or estimate a refund, credit, waiver, or timeline, and no saved reply used by this feature may contain one.
  • G2. Output classifier. Every draft passes a commitment detector (rules plus model) before display. If it fires, replace the offending sentence with a neutral placeholder, such as "[Agent: refund decision needed]", and tag the ticket.
  • G3. Send-time check. The detector also runs on the final edited text. If an agent's own text contains a refund commitment, Send requires an explicit confirmation ("This message commits to a refund"), which is logged. Agents may make commitments in their own words. The system may not.
  • G4. Audit. Retain drafts and final messages for the period Legal specifies. Support ops reviews a weekly sample of 100 refund-related tickets.
  • G5. Release bar. The detector must reach at least 98% recall on a labeled set of refund-commitment phrasings, including implicit ones like "we'll take care of that charge", before Billing goes live. Legal signs off on the test set.

6. Data and evaluation

Available: two years of resolved tickets with agent replies.

Data preparation (ML lead owns) - Use the queue where a ticket was resolved as the routing label, not where it first landed, since misroutes are the problem we're fixing. Also keep the initial queue to measure the historical misroute rate. - Remove PII before training or indexing. Get security review of the vendor and hosting setup if any external model is used. - Filter out replies that are outdated (old pricing, retired features, superseded policies). Support ops flags policy change dates so the pipeline can exclude earlier replies. - Mark replies that used a saved reply. This gives the saved-reply matching set and a baseline for the other 40%. - Have support ops label about 1,500 recent tickets for refund-commitment presence, routing, and draft quality. This set doubles as the gold evaluation set.

Evaluation - Split by time, not randomly: train on the earliest ~21 months, test on the most recent 3. Random splits will overstate performance because of seasonality and policy drift. - Routing: report accuracy and per-queue precision and recall, plus confusion between Billing and General, and Account access and Technical, which are likely weak spots. - Drafts: blind human review by senior agents on a 5-point rubric (accuracy, policy correctness, tone, completeness), with automatic policy-violation checks. Compare against the original agent reply. - Slice results by queue, topic, and saved-reply versus generated path.

7. Success metrics

MetricBaselineTarget
Mean first response time7h<2h at 90 days post-rollout (<4h at pilot)
Routing accuracy (first assignment)Measure in shadow≥92%
Drafts sent unedited or lightly editedn/a≥50% overall; ≥75% on saved-reply path
Draft discard raten/a<15%
Policy-incorrect drafts in weekly reviewn/a<2%
Refund commitments sent without confirmation flown/a0
Customer satisfaction on first-response ticketsCurrent CSATNo decline

First task: instrument the current split of first-response time into queue wait and agent handle time. Drafting mainly cuts handle time. If wait time dominates, we also need queue-ordering changes, such as surfacing the oldest tickets first, or the 2-hour target will be missed even with high draft quality. Throughput is about 43 tickets per agent per day, so handle time matters, but it is not necessarily the bottleneck.

8. Rollout

  1. Shadow mode (2–3 weeks). Run routing and drafting on live tickets without showing anything to agents. Compare with actual outcomes, set thresholds, and validate the refund detector.
  2. Pilot (3 weeks). Enable routing for all tickets. Show drafts to about 12 agents across General, Technical, and Account access.
  3. Expand. Enable drafts for all agents in those three queues once the pilot meets the quality targets.
  4. Billing. Enable only after G5 is met and Legal has signed off. Start with 6–8 agents, then expand.

Any queue can be switched off by a feature flag within minutes. Rollback triggers: any refund-commitment incident, a policy-incorrect rate above 5% in weekly review, or a CSAT drop of more than 3 points.

9. Risks and open questions

  • Automation bias. Agents may approve drafts without reading them. Mitigations: show sources, sample-audit sent messages, and track the edit rate by agent. A near-zero edit rate on generated-path drafts is a warning sign.
  • Stale content. Saved replies and help articles drift from policy. Support ops needs a named owner and a review cadence.
  • Language. We don't know the share of non-English tickets. Confirm before launch. If it is significant, scope it out or handle it separately.
  • Legal definition. Does "refund commitment" include credits, prorations, or fee waivers? Legal to define by the end of shadow mode, since it determines G1–G3.
  • Hosting. Confirm data-residency and vendor constraints for ticket content.
  • Threshold ownership. Decide who can change confidence thresholds in production (proposed: support-ops lead, with ML review).

10. Ownership

  • Engineering: integration with the ticketing system, sending-service enforcement (R10), logging, feature flags, latency.
  • ML: data preparation, routing model, drafting pipeline, refund detector, evaluation and monitoring.
  • Support ops: labeling, saved-reply and style-guide upkeep, pilot agents, weekly quality review, and thresholds.
  • Legal: the refund definition, the detector test set sign-off, and audit retention.

Grades and run details

Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly8%
  • passAddresses the actual decision78%
  • passRespects explicit constraints49%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims56%
  • passProduces the required deliverable74%
  • passSpecifies AI failure behaviour96%
  • passDefines evals and a launch bar99%
  • partialSuccess settles trade-offs78%
  • partialProposes tests that could fail64%
Run
Run
#1
API response time
46 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The PRD is complete and actionable for the named leads, requiring only the addition of explicit numeric thresholds for already-defined gates. No evidence or reasoning would need to be redone.

Output

PRD: AI First-Response Drafting and Ticket Routing

Status: Proposed Product: Ledgerly Support Audience: Engineering, ML, and Support Operations leads

1. Summary

Ledgerly receives about 9,000 support tickets per week. Tickets are handled by 42 agents across Billing, Technical, Account access, and General. Average first response time is seven hours, against a target of under two hours. About 60% of tickets receive one of roughly 40 saved replies.

This feature will recommend a queue and prepare a first-response draft when a new ticket arrives. An agent must review and approve every draft before it is sent. The feature is intended to reduce time spent triaging and composing routine responses, while leaving decisions and customer communication under agent control. It must never send a response autonomously or make a refund commitment on Ledgerly’s behalf.

2. Goals and non-goals

Goals - Reduce median time from ticket creation to first response, with an operational goal of under two hours. - Reduce agent effort on routine first responses by using appropriate saved replies and ticket-specific details. - Route tickets to the best-fit queue, while making uncertain recommendations visible and easy to correct. - Preserve agent review and control over every customer-facing response.

Non-goals - Automatically send, resolve, or close tickets. - Decide refund eligibility, promise a refund, or commit to refund timing. - Draft replies to later messages in an existing conversation in the initial release. - Replace queue ownership, escalation policies, or agents’ judgment.

3. Users and workflow

The primary user is a support agent reviewing new tickets in their assigned queue. Support Operations owns queue definitions, saved replies, and handling guidance. ML and Engineering own model quality, serving, integrations, and monitoring.

For each new ticket, the system will: 1. Read the ticket’s permitted content and available support context. 2. Recommend one of the four queues, with a confidence indicator and brief reason. 3. Generate a first-response draft, preferentially based on a relevant approved saved reply where appropriate. 4. Display the recommendation and draft in the agent workflow. 5. Let the agent edit, discard, or approve. Approval sends the response through the existing support system; no response is sent before approval.1 6. Record the final queue, draft disposition, edits, and outcome for evaluation.

Agents may change the queue before or after reviewing the draft. A draft must remain clearly marked as AI-generated until approved.

4. Functional requirements

Routing

  • Classify each ticket as Billing, Technical, Account access, or General.
  • Show the recommended queue, confidence, and a short explanation grounded in ticket content.
  • Allow agents to override the recommendation; the override becomes an evaluation signal, not an automatic training label.
  • Route low-confidence or ambiguous cases to General and flag them for review. Support Operations must set and approve the launch confidence threshold using offline and pilot results.
  • Do not infer urgency or bypass existing escalation rules in the initial release.

Drafting

  • Generate a concise, relevant first response based on the ticket and approved support materials available to the system.
  • Use a saved reply when it fits, adapting only with verified ticket details. Do not fabricate account status, actions taken, policy, or resolution.
  • When information is insufficient, ask a clear clarifying question or provide a safe acknowledgement rather than guessing.
  • Support agents can edit, discard, or approve the draft. No draft may be sent without an explicit agent action.
  • If the request concerns a refund, the draft must not promise, confirm, or imply a refund or refund timing. It should use approved noncommittal wording and leave eligibility and timing to an agent. Refund-related tickets should be visibly flagged for agent attention.

5. Data, model, and system requirements

Ledgerly has two years of resolved tickets and agent replies. ML should assess data coverage and quality before training: queue labels, duplicate or reopened cases, outdated answers, agent-specific wording, and tickets whose resolution depended on account information unavailable at ticket creation. Historical agent replies are examples, not policy; approved saved replies and current support guidance take precedence.

Create a time-based training, validation, and test split to reduce leakage from repeated tickets or changing policies. Evaluate routing and draft quality separately, including by queue, ticket type, and relevant risk category. Do not train on post-response information when evaluating first-response behavior. Remove or protect unnecessary sensitive data, and use only the customer and account context needed for support.

The serving path must retrieve only context the agent is authorized to see. Engineering must confirm integration points with the ticketing system, queue assignment, saved replies, and send/approval controls before implementation. Log model/version, recommendation, confidence, draft, agent edits, approval or discard, final queue, and response timestamps, subject to Ledgerly’s retention and access policies. Restrict access to these logs and avoid storing secrets or unnecessary personal data in model prompts.

6. Quality, safety, and success measures

Establish a pre-launch baseline and report weekly: - Median and 90th-percentile time to first response; proportion answered within two hours. - Queue recommendation accuracy and agent override rate, overall and by queue. - Draft approval rate, edit rate, discard rate, and time from ticket opening to send. - Agent-rated usefulness and sampled quality review results. - Refund-policy violations, fabricated claims, and other serious customer-impacting errors.

Suggested launch gates, to be confirmed by Support Operations and Legal: routing performance must be no worse than the current human baseline on the held-out set; pilot draft samples must meet an agreed quality rubric; and no refund commitment may appear in an approved draft during safety testing. Any refund commitment, unauthorized send, or material fabricated claim triggers immediate disablement of the affected capability and incident review. Do not optimize approval rate alone: agents must be able to reject drafts without penalty.

7. Rollout and operations

  1. Discovery and offline evaluation: Validate data, current queue definitions, approved saved replies, policy ownership, and integration feasibility. Build a labeled test set reviewed by Support Operations, with explicit refund and other high-risk examples.
  2. Shadow mode: Run recommendations without showing or applying them. Compare predicted queues and drafts with agent outcomes; inspect errors and subgroup performance.
  3. Limited pilot: Enable the workflow for a small, representative group of agents and ticket types. Keep agent approval mandatory. Support Operations reviews a daily sample and gathers structured agent feedback.
  4. Expansion: Expand to all 42 agents only after launch gates are met and Legal and Support Operations approve. Roll back by disabling drafting or routing independently if either degrades.

Support Operations owns the queue taxonomy, approved reply content, quality rubric, and agent training. ML owns evaluation, thresholds, drift monitoring, and model updates. Engineering owns access controls, reliable delivery, audit logs, feature flags, and rollback. A named on-call owner must be assigned before pilot launch. Review quality and performance weekly during rollout and after material model, policy, or saved-reply changes. Re-evaluate against a fresh sample regularly; do not promote agent edits to training data without review and policy checks.

8. Open decisions before build

Confirm the ticketing platform and integration surfaces; which customer/account fields are available at ticket creation; the authoritative source for current support policy; Legal’s approved refund language and review process; the minimum acceptable routing and draft-quality thresholds; and the retention period for prompts, drafts, and audit logs. These decisions are launch prerequisites, not reasons to relax agent approval or refund safeguards.

What a PM had to fix

  1. 1Test or gate too weakTighten the testTargeted repair

    What we’d changeAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly23%
  • passAddresses the actual decision66%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims39%
  • passProduces the required deliverable37%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar45%
  • partialSuccess settles trade-offs52%
  • partialProposes tests that could fail77%
Run
Run
#1
API response time
30 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongRightRight
Gemini 3.8 Flash · API

The PRD invents several current-state facts, including 7.1-hour FRT, manual triage details, channels, customer metadata, PII/CSAT/re-route data, and existing systems.

Sonnet 5.5 · API

All current-situation facts cited are from the supplied context or derived arithmetically, and causal statements are labelled as uncertain.

GPT-6 Luna · API

All current-state figures used come directly from the supplied context, with no invented facts about the current situation.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It likely exceeds the 1,000–1,500 word aim and its refund guardrail can still surface a draft containing refund language after stripping, so the prohibition is not cleanly enforced.

Sonnet 5.5 · API

It respects the legal no-auto-refund-commitment rule, requires agent approval, targets the named readers, and is within the word limit.

GPT-6 Luna · API

It respects the named audience, length, PRD form, and enforces the no-automated-refund-commitment constraint through mandatory approval, noncommittal refund wording, and flagging.

Avoids unsupported claimsWrongRightRight
Gemini 3.8 Flash · API

It presents several unestablished current-state assumptions as facts, such as existing channels, plan metadata, PII fields, CSAT data, and infrastructure.

Sonnet 5.5 · API

Hypotheses such as likely causes of delay are labelled as uncertain, and forecasts are tied to later measurement rather than asserted as fact.

GPT-6 Luna · API

No causes, forecasts, or interpretations are presented as established fact; uncertain items are framed as suggested, open, or to be confirmed.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a detailed PRD, but it is too long for the requested length and includes invented operational details that would need verification before use.

Sonnet 5.5 · API

The PRD is complete, practical, reader-appropriate and actionable for engineering, ML and support-ops leads.

GPT-6 Luna · API

The requested PRD is present, complete, written for engineering, ML, and support-operations leads, and within the requested length.

Success settles trade-offsWrongRightRight
Gemini 3.8 Flash · API

It sets metric targets but does not define a time frame for achieving FRT or an explicit rule for which goal wins when metrics conflict.

Sonnet 5.5 · API

It names target metrics with time frames and gives explicit trade-off rules such as no auto-send at any confidence and low-confidence routing to General.

GPT-6 Luna · API

It names a success metric with an under-two-hours target and weekly cadence, and gives trade-off rules such as disabling the affected capability on any refund commitment or unauthorized send.

All got wrong 1

Proposes tests that could failWrongWrongWrong
Gemini 3.8 Flash · API

Some gates have thresholds, but several lack clear measurement windows or explicit actions for failure, such as the GA FRT target and overall launch decision.

Sonnet 5.5 · API

Not every proposed gate has a numeric threshold, read-out window, and specified action; shadow and pilot go/no-go conditions are not fully quantified.

GPT-6 Luna · API

Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.

All got right 4

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits to building AssistRoute as a human-approved drafting and routing feature and frames the decision for engineering, ML, and support-operations leads.

Sonnet 5.5 · API

The output commits to a clear PRD design with routing, drafting, rollout and rollback conditions for the named leads.

GPT-6 Luna · API

The PRD commits unambiguously to AI-assisted queue recommendation and draft generation with mandatory agent approval, and specifies conditions for rollback or disablement.

Identifies material uncertaintyRightRightRight
Gemini 3.8 Flash · API

It identifies material unknowns such as routing accuracy, draft utilization, latency, model drift, and agent over-reliance, and proposes shadow mode, canary, and gates to resolve them.

Sonnet 5.5 · API

It explicitly lists open questions such as wait-vs-handle split, non-English share, legal definition of refund commitment, and hosting constraints, with owners and resolution paths.

GPT-6 Luna · API

It identifies material unknowns such as data quality, integration points, policy ownership, and confidence thresholds, and says how to resolve them through evaluation and piloting.

Specifies AI failure behaviourRightRightRight
Gemini 3.8 Flash · API

It specifies low-confidence routing to General with needs_triage, low-confidence drafting suppression, and refund-violation handling.

Sonnet 5.5 · API

It specifies low-confidence routing to General with triage, fallback clarifying drafts when content is insufficient, and draft-timeout behavior.

GPT-6 Luna · API

The PRD specifies low-confidence routing to General with flags, clarifying questions when information is insufficient, and immediate disablement for serious errors.

Defines evals and a launch barRightRightRight
Gemini 3.8 Flash · API

It defines offline validation, shadow mode, canary gates, routing accuracy, utilization, latency, and zero refund breach thresholds.

Sonnet 5.5 · API

It defines an evaluation set, blind review rubric, routing and draft metrics, refund detector recall, and launch bars.

GPT-6 Luna · API

It defines a pre-launch baseline, metrics, a held-out test set, launch gates, and a zero-refund-commitment safety condition.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD