Needs more than a quick edit8 of 10 checks passed
Unsupported claims about current workflow and data content would require re-checking evidence, and the data section's assumptions need validation before engineering can act.
Output
PRD: AI Draft & Route for Support Tickets
Sep 24, 2026 · @Martin Slaney
1. Summary and problem
Ledgerly's support team (42 agents, four queues: Billing, Technical, Account access, General) handles \~9,000 tickets a week. Average first response time (FRT) is 7 hours against a target of under 2.
Most of that delay is not writing time. It is triage and lookup1: \~60% of tickets are answered with one of \~40 saved replies, and 38% are billing. Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details.
This feature does two things when a ticket arrives:
- Routes it to the correct queue, with a confidence score.
- Drafts a first response, grounded in the saved-reply library, the customer's account data and similar past resolutions, for an agent to review.
Nothing is sent without an agent approving it. On top of that, any draft containing a refund commitment is hard-gated per Legal (§4.3).
2. Goals, non-goals and success metrics
Goals
- Bring average FRT under 2 hours without adding headcount.
- Cut misrouted tickets and the reassignment delay they cause.
- Reduce agent handling time on saved-reply tickets, so agents spend more time on the \~40% that need real investigation.
Non-goals (v1)
- Auto-sending any response, of any type.
- Drafting follow-ups beyond the first response.
- Taking actions in Ledgerly (issuing refunds, resetting passwords, changing plans). The model drafts text; agents act.
- Customer-facing chatbot or deflection.
Success metrics (measured per queue, against a 4-week pre-launch baseline)
| Metric | Target | Guardrail |
|---|---|---|
| Average FRT | < 2h | P90 FRT must not rise |
| Routing accuracy (final queue = predicted queue) | ≥ 92% overall, ≥ 90% per queue | Account access recall ≥ 95% |
| Draft acceptance (sent with light or no edits) | ≥ 50% of drafted tickets | — |
| Agent handling time on drafted tickets | −30% | — |
| CSAT on drafted tickets | No drop vs baseline (±1pt) | Reopen rate not up >1pt |
| Refund commitments sent without refund-approval step | 0 | Hard requirement |
"Light edit" is defined as a normalised edit distance below 0.2 between draft and sent text. ML owns the metric definition; Support Ops signs it off before pilot.
3. Users and core workflow
Users: support agents (reviewers of every draft); queue leads (monitor routing, handle overrides); Support Ops (owns saved replies, policies and the refund-approval rota).
Flow for a new ticket
- Ticket created (email or in-app form).
- Router predicts queue + confidence within 30s. High confidence → assigned to that queue. Low confidence → Triage view for a lead to assign in one click.
- Drafter produces a first response and attaches it to the ticket as an internal draft, with: the saved reply(s) it drew on, account facts it inserted, and any flags (refund, low confidence, missing data).
- Agent opens the ticket, reviews the draft, and chooses Send, Edit & send, or Discard (with a reason code).
- If the draft is refund-flagged, Send is replaced by Request refund approval (§4.3).
- Final queue, sent text and action are logged as training and evaluation signal.
4. Functional requirements
4.1 Routing
- R1. Classify every new ticket into Billing, Technical, Account access or General, with a calibrated confidence score.
- R2. Auto-assign when confidence ≥ threshold (set per queue from eval, starting target: ≥ 95% precision at that threshold). Below threshold → Triage view.
- R3. Account access is the costliest miss (locked-out customers).3 Tune for recall on this class; a ticket with any access signal and ambiguous classification goes to Account access, not General.
- R4. Agents can reassign in one click; every reassignment is logged with the original prediction.
- R5. Support Ops can switch routing to suggest-only per queue without a deploy.
4.2 Drafting
- D1. Generate a draft for every routed ticket in English within 60s of creation. Other languages: no draft in v1, flag only.
- D2. Retrieval first: identify the best-matching saved reply (or "none"). When one matches, the draft is that reply personalised with ticket and account context, not free text. When none matches, draft from similar resolved tickets and help-centre articles, and label it Free-form draft.
- D3. Account facts (plan, billing dates, invoice amounts, last payment status) come only from read-only lookups against Ledgerly's billing/account APIs, never from model memory. Any fact the model could not verify is left as a visible `[placeholder]` that blocks sending until filled.
- D4. Show sources inline: saved-reply ID, linked tickets/articles, API fields used.
- D5. Never promise timelines, credits, discounts, policy exceptions or refunds unless the saved reply itself contains them2 (refunds additionally gated, §4.3).
- D6. Skip drafting (flag only) for: legal threats, suspected fraud, data deletion/GDPR requests, security incidents, and abusive or distressed customers. Keyword + classifier; list owned by Support Ops.
4.3 Refund guardrail (Legal requirement)
Legal requires no automated sending of refund commitments. Because every draft already needs agent approval, we go further so a refund promise cannot slip through in a routine approve click:
- G1. A dedicated refund-commitment detector runs on every draft and on the final edited text at send time (agents may add refund language themselves). Tuned for recall ≥ 99% on a Legal-reviewed test set; false positives are acceptable.
- G2. If triggered, the ticket cannot be sent via one-click approval. The agent must confirm the refund is authorised under current policy (Support Ops defines who can authorise which amounts) via a separate confirmation step that is logged.
- G3. The drafter must never generate a refund commitment from free-form reasoning; refund language may only come from approved refund saved replies.
- G4. No bulk-approve action exists anywhere in the product.
- G5. Legal reviews the detector test set and the confirmation UX before pilot, and receives a monthly log of refund-flagged sends.
Open for Legal: does "automated sending" cover a one-click approve by an agent? This PRD assumes one-click approval is not automated but adds G2 as defence in depth. Confirm before build.
4.4 Agent experience
- A1. Draft appears pre-filled in the reply box, visibly marked as AI-drafted until edited.
- A2. Discard requires a reason: wrong answer, wrong tone, missing info, wrong queue, should not be drafted.
- A3. Agents never lose their normal tools; saved replies remain available manually.
- A4. Customers are not told a draft was AI-assisted (agent authors the sent message). Support Ops to confirm this against Ledgerly's AI disclosure policy.
5. ML approach, data and evaluation
Data. Two years of resolved tickets (\~900k) with final queue, agent replies and saved-reply usage.
- Label routing from the final queue, not the initial one; tickets that were reassigned are the most valuable examples.
- Map historic replies to saved-reply IDs where possible (exact/near-match), giving a supervised "which reply fits" dataset.
- Down-weight or exclude replies older than any policy or pricing change; Support Ops supplies the change dates. Stale answers are the main quality risk.
- Strip PII before any use outside the production data boundary; confirm data-processing terms with Legal/DPO for customer ticket content.
Approach. Start with the simplest thing that hits the targets: an LLM classifier (or fine-tuned small model if cost/latency requires) for routing; retrieval over saved replies + resolved tickets + help centre, then LLM personalisation for drafting; a separate refund detector. Model choice is ML's call, constrained by the latency targets in §3 and D1 and per-ticket cost.
Offline evaluation (gate to pilot)
- Held-out set: most recent 3 months, time-split (no leakage).
- Routing: per-queue precision/recall and calibration; auto-assign coverage at the chosen threshold.
- Drafting: saved-reply selection top-1 accuracy; 500-ticket human-graded sample scored by senior agents (correct, complete, on-policy, tone) with a ≥ 80% "send with light or no edit" bar.
- Refund detector: recall ≥ 99% on Legal-reviewed set.
- Factual accuracy: 0 invented account facts in the graded sample.
Online monitoring. Dashboards per queue for acceptance, edit distance, discard reasons, reassignments, FRT, CSAT and refund flags. Weekly review of a random sample of sent drafted replies by Support Ops QA.
6. Rollout, risks and open questions
Rollout
| Phase | Scope | Exit criteria |
|---|---|---|
| 0. Shadow (2 wks) | Route + draft on all tickets, invisible to agents | Offline bars met on live traffic |
| 1. Pilot (3 wks) | Billing queue, 8 agents; routing suggest-only | Acceptance ≥ 40%, no CSAT drop, 0 refund-guardrail misses |
| 2. Expand | All queues drafting; routing auto-assign above threshold | Metrics in §2 hold for 2 consecutive weeks |
| 3. Steady state | Monthly threshold review, quarterly saved-reply refresh | — |
Billing goes first: highest volume, most saved-reply coverage, and it exercises the refund guardrail early under close watch.
Risks
- Rubber-stamping. Agents approve drafts without reading them. Mitigate: placeholders block send, refund gate, QA sampling, track time-on-draft before send.
- Stale policy in drafts. Mitigate: saved replies as primary source, policy-change exclusion in training data, owner for saved-reply freshness.
- Misrouting Account access. Mitigate: recall-weighted threshold (R3), Triage fallback.
- Agent trust. Early bad drafts kill adoption. Mitigate: shadow phase, drafts only above a quality-confidence score; otherwise show the suggested saved reply only.
Open questions
- Legal: does agent one-click approval count as "automated sending" (§4.3)?
- Support Ops: who can authorise refunds, and at what amounts?
- Eng: which helpdesk platform APIs support pre-filled internal drafts and blocking send?
- DPO: approval to use historic ticket content with a model provider.
- Support Ops: AI-assistance disclosure to customers (A4).
Owners: Eng lead (integration, guardrail enforcement, UI), ML lead (models, eval, monitoring), Support Ops lead (saved replies, policy, pilot, QA).
What a PM had to fix
1Hypothesis stated as factReframe it as a hypothesisSubstantial rework
Source: Volumes
Average first response is 7 hours against a target of under 2.
What we’d changePresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.
2Constraint missedRestore the constraintTargeted repair
Source: Risk
Legal requires no automated sending of refund commitments.
What we’d changeState the refund rule once and consistently: generated drafts never contain refund commitments, and any commitment an agent adds goes through the separate confirmation step.
3Hypothesis stated as factReframe it as a hypothesisTargeted repair
What we’d changePresent it as an assumption to check with Support Ops. The brief doesn't rank the cost of misrouting by queue, and the routing design leans on it.
Claims the judge couldn’t find in the brief
- Most of that delay is not writing time. It is triage and lookup
- Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details
- Two years of resolved tickets (~900k) with final queue, agent replies and saved-reply usage are available
- Account access is the costliest miss (locked-out customers)
- Billing queue has the most saved-reply coverage
Grades and run details
Decision model 75 · LLM judge 8 of 11 checks
Decision model checks
- failUses the supplied evidence correctly48%
- passAddresses the actual decision73%
- passRespects explicit constraints34%
- passIdentifies material uncertainty85%
- failAvoids unsupported claims73%
- passProduces the required deliverable65%
- passSpecifies AI failure behaviour97%
- passDefines evals and a launch bar100%
- passSuccess settles trade-offs73%
- partialProposes tests that could fail68%
Artefacts
- link (link)
Run
- Run
- #1
- Time to output
- 2.0 min
- Submitted
- 24 Sept 2026