Needs more than a quick edit7 of 10 checks passed
It is largely usable but invents current systems not in the brief, so a PM would need to verify the evidence before acting. It also lacks an explicit trade-off rule and complete test windows/actions, requiring more than a quick edit.
Output
PRD: AI-Assisted First Response and Ticket Routing
Product: Ledgerly Support Assist Audience: Engineering, ML and Support Operations leads Status: Proposed MVP; thresholds below are launch targets, not measured results Accountable owner: Support Operations lead, with Engineering and ML leads owning delivery and model quality
1. Problem and objective
Ledgerly’s 42 support agents handle 9,000 tickets per week across Billing, Technical, Account access and General. Billing represents 38% of volume—approximately 3,420 tickets weekly. Average first-response time is seven hours, against a target of under two hours. Approximately 60% of tickets receive one of roughly 40 saved replies.
The feature will classify new tickets into the appropriate queue and prepare a grounded first-response draft for an agent to review. It should reduce queue-selection work and repetitive writing without delegating customer communication or financial commitments to the model.
Every outbound response requires explicit agent approval. Legal prohibits automated sending of refund commitments. The MVP will have no automated-send path and will not generate refund promises; agents must author any commitment using the existing authorized refund process.
2. Goals and success measures
Primary goals:
- Reduce average time to the first substantive human-approved response to under two hours during a staffed pilot.
- Reduce median active agent time spent preparing first responses by at least 30%.
- Route tickets accurately while avoiding hidden misroutes.
- Preserve response correctness, account security and customer satisfaction.
Measure first-response time from ticket creation to the first substantive response sent by an agent. Automated receipts do not count. Report calendar-hour and staffed-hour results separately, plus median and p90, to prevent averages hiding long waits.
Additional proposed launch and pilot thresholds:
| Measure | Target |
|---|---|
| Queue-classification accuracy | ≥95% on the adjudicated holdout |
| Account access routing recall | ≥98% |
| High-confidence automatic-routing precision | ≥98% |
| Drafts usable with no substantive correction | ≥80% in blinded review |
| Unsupported financial or security commitments | Zero observed in launch evaluation |
| Draft availability latency | p95 ≤15 seconds after ingestion |
| Pilot quality | No material deterioration in QA scores or customer satisfaction |
Zero observed errors is a release gate, not proof of zero production risk. ML must report sample sizes and confidence intervals. Support Operations must confirm staffing coverage: drafting improvements alone cannot guarantee the response-time target.
3. Scope
MVP includes:
- Classification into Billing, Technical, Account access or General.
- Automatic routing only above validated, queue-specific confidence thresholds.
- Retrieval of approved saved replies and current support documentation.
- A first-response draft with internal source references, risk flags and suggested clarifying questions.
- Agent controls to edit, discard, regenerate, change queue and approve/send.
- Audit logs, monitoring and immediate disable controls.
Excluded: automatic sending, subsequent-turn assistance, ticket resolution, refunds, account changes, security verification decisions and customer-facing AI chat.
Ticket channels, languages and attachment formats are not specified. Support Operations will inventory them before implementation. MVP supports only validated text channels and languages; unsupported inputs receive normal human triage without a generated draft.
4. User workflow and functional requirements
- Ingest: On creation of a new eligible ticket, capture its text, subject and approved metadata. Use only authorized customer/account context available to the assigned support role.
- Assess: Predict a queue, calibrated confidence and risk flags such as refund request, account compromise or insufficient information.
- Route: Assign high-confidence tickets to the predicted queue. Send uncertain cases to the existing manual-triage destination, provisionally General. Surface uncertainty prominently; do not treat General as a confident classification.
- Draft: Retrieve relevant approved material and produce a concise response addressing the request. Prefer adapting an applicable saved reply over open-ended generation.
- Review: Show the draft, suggested queue, risk flags and source links in the agent workspace. Clearly label the text “AI draft—not sent.”
- Approve/send: The existing send action requires an explicit authenticated agent action. No background job, timeout or model output may trigger sending.
- Learn: Record queue corrections, edits, rejection reasons and quality reviews for controlled evaluation and future retraining.
Drafts must not invent transactions, troubleshooting outcomes, refund eligibility or account status. Where evidence is missing, ask for necessary information or acknowledge that an agent must investigate. Avoid requesting passwords, full payment-card details or other unnecessary sensitive data.
Refund-related drafts may acknowledge the request and explain approved next steps, but must not promise an amount, eligibility or processing date. Flag these tickets for Billing review. Agents may add commitments only after authorized verification.
A draft becomes stale if relevant ticket content or account context changes. Disable approval until it is refreshed or explicitly reviewed against the updated context. Manual queue changes take precedence over later model results.
5. Data and ML approach
Two years of resolved tickets and agent replies are available. Historical replies are examples, not authoritative policy: outdated guidance and unauthorized commitments must not be reproduced.
ML and Support Operations will:
- Define queue labels and rules for mixed-intent tickets. Account-security concerns take precedence over routine billing questions; Support Operations must approve the complete precedence matrix.
- Audit final queue labels and sample ambiguous cases for expert adjudication.
- Remove duplicates, signatures and irrelevant quoted history; redact unnecessary personal and payment information.
- Preserve ticket/thread/customer grouping across splits to reduce leakage.
- Use a chronological training, validation and held-out test split, with the newest period reserved for testing.
- Evaluate only information available when the ticket arrived. Later replies and resolutions may provide labels but must never become runtime inputs.
Benchmark saved-reply retrieval and a conventional classifier before introducing more complex models. Select the simplest approach meeting quality, latency and operational requirements.
Generation will use current, versioned approved content. Historical replies may support offline training subject to privacy and quality approval, but must not serve as an unrestricted runtime knowledge source. Conflicting or missing sources trigger a clarification or human-investigation draft rather than a guessed answer.
Customer text and retrieved content are untrusted inputs. Instructions embedded in tickets must not change system rules, authorize actions or bypass review.
6. Engineering design and controls
Implement an asynchronous pipeline behind feature flags:
Ticket event → eligibility check → classifier → routing decision → retrieval → generation → policy validation → draft storage → agent UI.
Persist ticket ID/version, model and prompt versions, retrieved document versions, queue scores, draft state, risk flags and agent actions. Use idempotency keys to prevent duplicate processing and concurrency controls to avoid overwriting agent work.
The AI service must have no credential or permission to send messages, issue refunds or modify accounts. Routing permissions must be limited to approved queues. Outbound messages remain controlled by the existing authenticated support application.
Post-generation checks will block prohibited promises and unsupported sensitive claims. Failed checks suppress the draft and display a reason; they must not merely append a disclaimer.
Apply role-based access, encryption and existing retention rules. Any external model provider requires Security and Legal approval, including contractual restrictions on retention and training use. Log identifiers and operational metadata where possible, not unrestricted ticket bodies.
On timeout, provider outage or validation failure, preserve normal ticket handling. Keep the ticket visible, apply manual triage where necessary and show “Draft unavailable.” Never delay ticket intake while waiting for AI.
7. Evaluation and release gates
Build an adjudicated test set covering all queues, common saved-reply cases, mixed intent, sparse descriptions, refund disputes, account compromise, outdated-policy examples and prompt-injection attempts. Report overall and per-queue performance; oversampled risk cases must also be reported separately from production-weighted results.
Two support reviewers will score drafts for correctness, relevance, completeness, tone and policy compliance, with disagreements adjudicated. Distinguish cosmetic edits from substantive corrections. Track source support and appropriate abstention, not acceptance rate alone.
Required acceptance tests include:
- A refund-request draft contains no commitment.
- No model or pipeline component can invoke sending.
- Unsupported facts cause omission or escalation.
- Ticket updates invalidate stale drafts.
- Human rerouting is not overwritten.
- Provider failure leaves tickets accessible.
- Duplicate events create neither duplicate drafts nor duplicate sends.
- Agent permissions also constrain retrieved context.
Production expansion requires meeting quality thresholds and passing all safety-critical tests. Any unauthorized financial promise in evaluation blocks launch pending remediation.
8. Rollout, ownership and operations
Indicative eight-week plan:
- Weeks 1–2: Support Operations approves taxonomy, saved replies, triage rules and baseline measurements; ML audits data; Engineering completes integration and security design.
- Weeks 3–4: Build the pipeline, UI, evaluation harness and operational controls.
- Weeks 5–6: Run offline evaluation and shadow mode without changing queues or showing drafts.
- Weeks 7–8: Pilot with selected agents, then expand by queue only after gates pass.
Use a randomized holdout where practical, balanced by queue and shift, to compare handling time, response time, corrections and customer outcomes.
Engineering owns availability, access controls and rollback. ML owns evaluation, calibration and drift monitoring. Support Operations owns policy content, QA, agent training and staffing. Legal approves refund wording and provider terms.
Provide separate kill switches for routing and drafting. Suspected unauthorized commitments, data exposure or account-security mishandling trigger immediate suspension of the affected capability and incident review. Review corrections and rejected drafts weekly; retraining requires fresh holdout evaluation, not automatic ingestion of agent edits.
Check by check
Got wrong · 2
- Uses the supplied evidence correctlyIt invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
- Proposes tests that could failSeveral acceptance tests and kill criteria lack numeric thresholds, measurement windows, or the action each outcome triggers.
Mixed · 1
- Success settles trade-offsIt does not give an explicit trade-off rule such as accepting lower coverage to preserve precision above a stated level.The two graders disagreed on this one.
Got right · 7
- Addresses the actual decisionIt commits to an MVP with a clear approval flow, release gates, and conditions that would change rollout.
- Respects explicit constraintsIt respects the named audience, length, human approval requirement, and legal prohibition on automated refund commitments.
- Identifies material uncertaintyIt names missing channels/languages, unmeasured thresholds, sample-size limitations, and how evaluation and pilot results would resolve them.
- Avoids unsupported claimsNormative requirements and design rules are framed as proposals, not as established facts from the supplied evidence.
- Produces the required deliverableIt is a complete PRD for engineering, ML, and support-operations leads that could start work from.
- Specifies AI failure behaviourIt specifies low-confidence routing to manual triage, abstention or clarification when evidence is missing, stale-draft disabling, and outage fallback.
- Defines evals and a launch barIt defines an adjudicated test set, quality metrics, thresholds, and launch gates.
Claims the judge couldn’t find in the brief
- The existing send action requires an explicit authenticated agent action.
- There is an existing manual-triage destination, provisionally General.
- There is an existing authorized refund process agents must use.
Grades and run details
Decision model 90 · LLM judge 7 of 11 checks
Decision model checks
- partialUses the supplied evidence correctly7%
- passAddresses the actual decision90%
- passRespects explicit constraints39%
- passIdentifies material uncertainty44%
- passAvoids unsupported claims49%
- passProduces the required deliverable75%
- passSpecifies AI failure behaviour100%
- passDefines evals and a launch bar100%
- passSuccess settles trade-offs26%
- partialProposes tests that could fail69%
Run
- Run
- #1
- API response time
- 49 s
- Submitted
- 29 Sept 2026