Usable with a quick edit9 of 10 checks passed
The PRD is thorough, actionable, and requires at most minor wording edits before engineering, ML, and support-ops leads can start work.
Output
Ledgerly: AI first-response drafts and ticket routing
Status: Proposed MVP requirements. Numerical gates below are proposed targets, not observed performance. Product owns scope; Engineering owns workflow integrity; ML owns model quality; Support Operations owns queue policy, content and review quality.
1. Problem and outcome
Ledgerly’s 42 support agents handle 9,000 tickets weekly across Billing, Technical, Account access and General. Billing represents 38%, approximately 3,420 tickets weekly. Average first response is seven hours against a target below two. Approximately 60% of tickets use one of 40 saved replies.
The feature should reduce triage and composition time without shifting effort into correcting poor drafts or recovering misrouted tickets. Every outgoing draft requires human approval. Legal prohibits automated sending of refund commitments.
Faster drafting alone may not close the five-hour gap. Before piloting, Support Operations must measure arrival-to-assignment, assignment-to-review and review-to-send delays by queue and staffed hours, then identify any coverage or staffing changes needed.
2. Scope and operating boundary
MVP processes newly created English-language, text-based tickets in the existing support workspace. It assigns a queue and prepares a first-response draft. Threads, ticket metadata and approved support content are inputs. Attachments are not interpreted; attachment-dependent or unsupported-language tickets remain available for manual handling.
Excluded: follow-up generation, autonomous sending, refunds, account changes, payment actions and staffing optimisation. Existing spam, security and priority rules run first and cannot be overridden by the model.2
Routing and drafting operate independently. A drafting failure must not prevent routing or agent access. An uncertain route must not prevent a useful draft.
3. Agent workflow
- Ticket creation starts asynchronous processing. The ticket is immediately visible and its response clock starts at original receipt.
- The system records a route, reason and confidence tier, then prepares a draft where supported.
- The agent sees the original message, assigned queue, editable draft, source references and warnings. Sources and warnings are internal only.
- The agent can change queue, edit, regenerate, discard or write manually. Regeneration never overwrites unsaved edits without confirmation.
- Selecting Approve and send authorises the exact visible text. There is no bulk approval.
Draft states are pending, ready, unavailable, stale and sent. Queue changes preserve receipt time and draft history. New customer messages or changes to relevant ticket context mark drafts stale and require renewed review. Show actionable failure messages and keep the manual composer available.
4. Routing requirements
Support Operations owns this initial taxonomy:
| Queue | Primary issue |
|---|---|
| Billing | Charges, subscriptions, invoices, cancellations and refund requests |
| Technical | Errors, integrations, imports and malfunctioning features |
| Account access | Login, authentication, permissions and suspected account takeover |
| General | Product guidance and genuinely uncategorised enquiries |
For multiple intents, suspected account compromise takes precedence, followed by access-blocking issues; otherwise route by the customer’s main requested resolution. Record secondary intents for the receiving agent.
Automatically assign only when a queue-specific threshold meets the evaluation gate below. Confidence must be calibrated against labelled examples, not taken from a model’s self-reported certainty.
Below threshold, assign to General with a distinct Needs triage status and show the leading suggestions. This is a fallback assignment, not a successful classification. Support Operations assigns a named triage owner each shift and reviews these tickets at least every 30 minutes during staffed hours. Existing out-of-hours escalation remains in force.
Agents can override any route and optionally record a reason. Never automatically reroute after an agent takes ownership. Log overrides for review, not immediate retraining.
5. Draft content and refund controls
Start with retrieval from the approximately 40 saved replies and current, Support Operations-approved help and policy content. Adapt an applicable reply before attempting a novel answer. Each source has an owner, version and review date; withdrawn content becomes unavailable immediately.
Drafts must answer the stated question, request necessary missing details and cite supporting sources internally. They must not invent account facts, troubleshooting outcomes, eligibility, amounts or dates. Account-specific claims require authorised, current account context. Without it, draft a clarification or indicate that an agent must investigate.
For refund requests, MVP drafts may acknowledge the request and explain approved review steps, but must not promise eligibility, payment amounts or payment dates. Agents may manually add commitments only under Ledgerly’s existing refund authority policy. Flag refund-related tickets and require an explicit acknowledgement when the final text contains a detected commitment. This check assists reviewers; it is not the legal enforcement boundary.
The enforcement boundary is the sending service: generation workers have no send credentials. Every AI-assisted send requires an authenticated agent approval bound to ticket ID, recipient and exact message version. Any subsequent edit invalidates approval. Retries cannot send duplicate messages. Existing automation must not consume AI drafts as sendable replies. Therefore, missed refund detection cannot trigger autonomous sending.
6. Data and ML approach
Use the two-year resolved-ticket archive to learn routing patterns and evaluate drafting. Historical agent replies are examples, not authoritative policy. Resolution does not establish correctness or refund permission.
Support Operations must relabel a representative sample against today’s taxonomy, with two reviewers and adjudication for disagreements. Redact credentials, payment details and unnecessary personal information. Preserve tenant boundaries and restrict access to approved training personnel and services.
Split chronologically into training, validation and a locked recent test set. Keep complete conversations and duplicate ticket clusters in one split. At inference and evaluation, expose only information available before the first response; later replies and final queue labels must not leak into inputs.
Begin with a classifier plus retrieval-grounded generation. Fine-tuning is optional and requires measurable improvement over this baseline. No automatic learning from approvals. Provider data use and retention must be approved before production data leaves Ledgerly.
Treat customer text and retrieved content as untrusted data: embedded instructions cannot change policies, access other tenants or invoke privileged actions.
7. Evaluation and launch gates
Build a locked test set of at least 1,000 tickets, stratified across queues, with at least 150 per queue. Include saved-reply matches, ambiguous intents, refunds, outdated policies, missing context, prompt injection and sensitive-data cases. Report volume-weighted results and per-queue results separately.
Required gates:
- Routing: at least 95% precision among automatically assigned tickets in each queue. Report confidence intervals, recall, coverage and the confusion matrix. Target at least 60% overall automatic-routing coverage; do not lower precision to achieve coverage.
- Drafts: at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.1 Score factual accuracy, policy compliance, completeness and tone separately.
- Critical errors: zero observed cross-tenant disclosures, unsupported refund commitments in generated drafts or account-security instructions that bypass verification. Any occurrence blocks release pending correction and regression testing.
- Workflow: all permission, stale-approval, recipient-change, duplicate-event and retry tests pass. Attempts to send without valid approval must fail.
Use blinded human review and adjudicate disagreements. Automated grading can assist sampling but cannot determine safety gates alone. Zero observed errors is a test result, not proof of zero production risk.
8. Pilot, measurement and operations
Run one week in shadow mode across all queues, followed by a two-week pilot with agents from every queue. Randomise eligible tickets within queue and shift between assisted and existing workflows; account for shared-agent effects when interpreting results.
Primary pilot gate: at least 20% lower mean receipt-to-first-human-response time versus control. Track progress towards the under-two-hour target across all incoming tickets, including unsupported and fallback cases. Also report median, p90, percentage answered within two hours, review time, queue transfers, correction severity, backlog age and customer satisfaction. Draft creation is not a first response.
Expand only if offline gates hold, human handling time improves and transfer rates and customer satisfaction show no material deterioration. Define tolerances and adequate sample sizes before the pilot; extend measurement when inconclusive.
Engineering provides independent routing and drafting kill switches, audit logs and alerts. Target p95 draft readiness within 60 seconds; after timeout, mark unavailable and retain manual handling. Retries must not overwrite agent work.
Support Operations samples 50 assisted tickets weekly, oversampling refunds and overrides. ML monitors quality and drift by queue and content version. Any unauthorised send or cross-tenant disclosure triggers immediate suspension of the affected feature and incident review.
Before pilot launch, Engineering verifies send-path enforcement and integration contracts; ML publishes evaluation results; Support Operations approves taxonomy, sources and shift ownership; Legal confirms the refund workflow.
What a PM had to fix
1Test or gate too weakTighten the testQuick edit
What we’d changeSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.
2Invented evidenceVerify or remove the claimQuick edit
What we’d changeThe brief doesn't mention existing spam, security or priority rules, or out-of-hours escalation. Confirm they exist before the design relies on them.
Check by check
Mixed · 1
- Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the supplied context or arithmetic, with no invention.The two graders disagreed on this one.
Got right · 9
- Addresses the actual decisionThe PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
- Respects explicit constraintsThe no-automated-refunds requirement is enforced via draft rules, approval checks, and a send-path enforcement boundary; length and reader constraints are met.
- Identifies material uncertaintyUnknowns like actual delay components, zero-error test limits, and pilot inconclusiveness are named, with resolution steps specified.
- Avoids unsupported claimsHypotheses (e.g., 'faster drafting alone may not close the gap') are clearly flagged as possibilities, not fact.
- Produces the required deliverableA complete PRD with scope, workflow, ML approach, evals, and pilot plan is delivered for the target leads, within ~1,200 words.
- Specifies AI failure behaviourFallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
- Defines evals and a launch barA 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
- Success settles trade-offsSuccess metric (≥20% reduction in response time) with a pilot time frame, and an explicit precision-over-coverage trade-off rule are given.
- Proposes tests that could failAll evaluation and pilot gates have numeric thresholds, a two-week measurement window, and defined actions (block release, suspend feature).
Grades and run details
Decision model 85 · LLM judge 11 of 11 checks
Decision model checks
- failUses the supplied evidence correctly2%
- passAddresses the actual decision85%
- passRespects explicit constraints50%
- passIdentifies material uncertainty83%
- passAvoids unsupported claims75%
- passProduces the required deliverable75%
- passSpecifies AI failure behaviour100%
- passDefines evals and a launch bar100%
- passSuccess settles trade-offs89%
- partialProposes tests that could fail73%
Run
- Run
- #1
- Time to output
- 1.8 min
- Submitted
- 24 Sept 2026