Usable with a quick edit9 of 10 checks passed
The PRD is usable as-is with only small edits to quantify the pilot go/no-go thresholds; the evidence and reasoning are sound and would not need redoing.
Output
PRD: AI-Drafted First Responses and Ticket Routing
Status: Draft for build planning Audience: Engineering, ML, and Support Operations leads Product: Ledgerly Support
1. Problem
Ledgerly receives about 9,000 support tickets a week, handled by 42 agents across four queues: Billing, Technical, Account access, and General. Billing is 38% of volume (about 3,400 tickets a week). Average first response is 7 hours against a target of under 2.
Two things likely drive the delay, and we don't yet know their split:
- Wait time. Tickets sit in a queue, or in the wrong queue, before anyone opens them.
- Handle time. Agents spend time writing or assembling a first reply. About 60% of tickets are already answered with one of roughly 40 saved replies, so much of this work is repetitive.
This feature has two parts. It routes each new ticket to the right queue, and it drafts a first response for the agent to review, edit, and send. Nothing is ever sent without an agent's approval.
2. Goals and non-goals
Goals 1. Bring mean first response time under 2 hours within 90 days of full rollout. 2. Route at least 92% of tickets to the correct queue on first assignment. 3. Have agents send at least 50% of drafts with no or light edits. 4. Ensure no draft or sent message contains an unreviewed refund commitment (see §5).
Non-goals - Auto-sending any message, in any queue, at any confidence level. - Resolving tickets, or taking actions such as issuing refunds, changing plans, or resetting access. - Replacing follow-up replies after the first response (later phase). - Changing queue structure or staffing.
3. Users
- Support agents (42): review, edit, approve, or discard drafts, and correct routing.
- Support-ops leads: own saved replies, policy content, thresholds, and quality review.
- Legal/compliance: consulted on the refund guardrail and audit logs.
4. Functional requirements
4.1 Routing
- R1. On ticket creation, classify into one of the four queues and assign it automatically.
- R2. Store a confidence score with each assignment. Below a threshold set by support ops (initially tuned so about 10% of tickets fall below it), assign to General with a "needs triage" tag and show the top two suggested queues.
- R3. Agents can re-route in one click and give an optional reason. Every re-route is logged as a labeled correction.
- R4. Also emit tags used by drafting: a topic label (for example, invoice question, failed payment, login lockout) and a refund-related flag.
4.2 Draft generation
- R5. A draft must be ready in the agent's ticket view within 60 seconds of ticket creation, and never block the ticket from appearing. If drafting fails or times out, the ticket appears as it does today.
- R6. Use two drafting paths:
- Saved-reply path. If a ticket matches one of the ~40 saved replies above a similarity threshold, use that reply with variables filled in (name, plan, invoice number). The wording stays as approved by support ops.
- Generated path. Otherwise, generate a draft grounded in retrieved help-center articles and similar past resolved tickets. Every factual claim about policy, pricing, or product behavior must come from retrieved content. If the content is insufficient, the draft asks the customer a clarifying question or says an agent is investigating, rather than guessing.
- R7. Each draft shows which path produced it and its sources (saved reply ID or article links) so agents can verify quickly.
- R8. Match the tone and format of the best-rated historical agent replies. Support ops maintains the style guide.
4.3 Agent workflow
- R9. Drafts appear pre-filled in the reply composer with three actions: Send (after optional edits), Discard, and Regenerate. Discarding takes an optional reason from a short list.
- R10. The system has no auto-send code path. Sending requires an authenticated agent action, and this is enforced in the sending service, not just the UI.
- R11. Log for each ticket: draft text, final sent text, edit distance, agent ID, time from ticket open to send, and discard or re-route reasons.
4.4 Account-access safeguards
- R12. Drafts for Account access tickets must never confirm whether an account exists, reveal account details, or state that access was restored or changed. They may give standard verification instructions from approved content only.
5. Refund guardrail (legal requirement)
Legal requires no automated sending of refund commitments. Because agents approve every draft, the primary control is R10. We add layered controls so the model never puts a commitment in front of an agent as if it were approved policy.
- G1. Prompt and template rules. Drafts may acknowledge a refund request and say it is being reviewed. They must not promise, imply, or estimate a refund, credit, waiver, or timeline, and no saved reply used by this feature may contain one.
- G2. Output classifier. Every draft passes a commitment detector (rules plus model) before display. If it fires, replace the offending sentence with a neutral placeholder, such as "[Agent: refund decision needed]", and tag the ticket.
- G3. Send-time check. The detector also runs on the final edited text. If an agent's own text contains a refund commitment, Send requires an explicit confirmation ("This message commits to a refund"), which is logged. Agents may make commitments in their own words. The system may not.
- G4. Audit. Retain drafts and final messages for the period Legal specifies. Support ops reviews a weekly sample of 100 refund-related tickets.
- G5. Release bar. The detector must reach at least 98% recall on a labeled set of refund-commitment phrasings, including implicit ones like "we'll take care of that charge", before Billing goes live. Legal signs off on the test set.
6. Data and evaluation
Available: two years of resolved tickets with agent replies.
Data preparation (ML lead owns) - Use the queue where a ticket was resolved as the routing label, not where it first landed, since misroutes are the problem we're fixing. Also keep the initial queue to measure the historical misroute rate. - Remove PII before training or indexing. Get security review of the vendor and hosting setup if any external model is used. - Filter out replies that are outdated (old pricing, retired features, superseded policies). Support ops flags policy change dates so the pipeline can exclude earlier replies. - Mark replies that used a saved reply. This gives the saved-reply matching set and a baseline for the other 40%. - Have support ops label about 1,500 recent tickets for refund-commitment presence, routing, and draft quality. This set doubles as the gold evaluation set.
Evaluation - Split by time, not randomly: train on the earliest ~21 months, test on the most recent 3. Random splits will overstate performance because of seasonality and policy drift. - Routing: report accuracy and per-queue precision and recall, plus confusion between Billing and General, and Account access and Technical, which are likely weak spots. - Drafts: blind human review by senior agents on a 5-point rubric (accuracy, policy correctness, tone, completeness), with automatic policy-violation checks. Compare against the original agent reply. - Slice results by queue, topic, and saved-reply versus generated path.
7. Success metrics
| Metric | Baseline | Target |
|---|---|---|
| Mean first response time | 7h | <2h at 90 days post-rollout (<4h at pilot) |
| Routing accuracy (first assignment) | Measure in shadow | ≥92% |
| Drafts sent unedited or lightly edited | n/a | ≥50% overall; ≥75% on saved-reply path |
| Draft discard rate | n/a | <15% |
| Policy-incorrect drafts in weekly review | n/a | <2% |
| Refund commitments sent without confirmation flow | n/a | 0 |
| Customer satisfaction on first-response tickets | Current CSAT | No decline |
First task: instrument the current split of first-response time into queue wait and agent handle time. Drafting mainly cuts handle time. If wait time dominates, we also need queue-ordering changes, such as surfacing the oldest tickets first, or the 2-hour target will be missed even with high draft quality. Throughput is about 43 tickets per agent per day, so handle time matters, but it is not necessarily the bottleneck.
8. Rollout
- Shadow mode (2–3 weeks). Run routing and drafting on live tickets without showing anything to agents. Compare with actual outcomes, set thresholds, and validate the refund detector.
- Pilot (3 weeks). Enable routing for all tickets. Show drafts to about 12 agents across General, Technical, and Account access.
- Expand. Enable drafts for all agents in those three queues once the pilot meets the quality targets.
- Billing. Enable only after G5 is met and Legal has signed off. Start with 6–8 agents, then expand.
Any queue can be switched off by a feature flag within minutes. Rollback triggers: any refund-commitment incident, a policy-incorrect rate above 5% in weekly review, or a CSAT drop of more than 3 points.
9. Risks and open questions
- Automation bias. Agents may approve drafts without reading them. Mitigations: show sources, sample-audit sent messages, and track the edit rate by agent. A near-zero edit rate on generated-path drafts is a warning sign.
- Stale content. Saved replies and help articles drift from policy. Support ops needs a named owner and a review cadence.
- Language. We don't know the share of non-English tickets. Confirm before launch. If it is significant, scope it out or handle it separately.
- Legal definition. Does "refund commitment" include credits, prorations, or fee waivers? Legal to define by the end of shadow mode, since it determines G1–G3.
- Hosting. Confirm data-residency and vendor constraints for ticket content.
- Threshold ownership. Decide who can change confidence thresholds in production (proposed: support-ops lead, with ML review).
10. Ownership
- Engineering: integration with the ticketing system, sending-service enforcement (R10), logging, feature flags, latency.
- ML: data preparation, routing model, drafting pipeline, refund detector, evaluation and monitoring.
- Support ops: labeling, saved-reply and style-guide upkeep, pilot agents, weekly quality review, and thresholds.
- Legal: the refund definition, the detector test set sign-off, and audit retention.
Check by check
Got wrong · 1
- Proposes tests that could failNot every proposed gate has a numeric threshold, read-out window, and specified action; shadow and pilot go/no-go conditions are not fully quantified.
Got right · 9
- Uses the supplied evidence correctlyAll current-situation facts cited are from the supplied context or derived arithmetically, and causal statements are labelled as uncertain.
- Addresses the actual decisionThe output commits to a clear PRD design with routing, drafting, rollout and rollback conditions for the named leads.
- Respects explicit constraintsIt respects the legal no-auto-refund-commitment rule, requires agent approval, targets the named readers, and is within the word limit.
- Identifies material uncertaintyIt explicitly lists open questions such as wait-vs-handle split, non-English share, legal definition of refund commitment, and hosting constraints, with owners and resolution paths.
- Avoids unsupported claimsHypotheses such as likely causes of delay are labelled as uncertain, and forecasts are tied to later measurement rather than asserted as fact.
- Produces the required deliverableThe PRD is complete, practical, reader-appropriate and actionable for engineering, ML and support-ops leads.
- Specifies AI failure behaviourIt specifies low-confidence routing to General with triage, fallback clarifying drafts when content is insufficient, and draft-timeout behavior.
- Defines evals and a launch barIt defines an evaluation set, blind review rubric, routing and draft metrics, refund detector recall, and launch bars.
- Success settles trade-offsIt names target metrics with time frames and gives explicit trade-off rules such as no auto-send at any confidence and low-confidence routing to General.
Grades and run details
Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
- partialUses the supplied evidence correctly9%
- passAddresses the actual decision77%
- passRespects explicit constraints50%
- passIdentifies material uncertainty98%
- passAvoids unsupported claims57%
- passProduces the required deliverable73%
- passSpecifies AI failure behaviour97%
- passDefines evals and a launch bar99%
- partialSuccess settles trade-offs63%
- partialProposes tests that could fail65%
Run
- Run
- #1
- API response time
- 46 s
- Submitted
- 29 Sept 2026