Needs more than a quick edit4 of 10 checks passed
Contains an unsupported current routing accuracy figure, missing evaluation plan/launch bar, and no trade‑off rules; a PM would need to re‑check evidence and add substantial missing sections.
Output
Product Requirements Document (PRD): AI-Powered Ticket Triage and Response Draft Generation ("AutoDraft & Route")
---
1. Document Overview
| Metadata | Details |
|---|---|
| :--- | :--- |
| Project Name | AutoDraft & Route (Ledgerly Support Intelligence) |
| Target Audience | Engineering, Machine Learning (ML), and Support Operations (Ops) Leads |
| Document Status | Draft / Ready for Engineering Review |
| Target Launch | Q1 2027 |
| Goal | Reduce Average First Response Time (FRT) from 7 hours to <2 hours while maintaining human-in-the-loop safety and strict legal compliance on financial commitments. |
---
2. Background & Problem Statement
2.1 Current State & Metrics
Ledgerly is a premier accounting software platform designed specifically for small businesses. As our customer base scales, our support infrastructure is facing severe volume pressures: * Support Team Size: 42 agents distributed across four specialized queues: Billing, Technical, Account Access, and General. * Ticket Volume: ~9,000 incoming tickets per week, with 38% (~3,420 tickets/week) falling into the Billing queue. * First Response Time (FRT): Current average is 7 hours, severely lagging behind our target of under 2 hours. * Repetitive Workflows: Approximately 60% of all incoming tickets are resolved using one of roughly 40 standard saved replies (e.g., password resets, invoice retrieval, subscription tier changes).
2.2 The Opportunity
We possess two years of historical resolved tickets, including full conversation histories and agent replies. This corpus serves as a robust training and evaluation dataset for an ML-driven system that can: 1. Accurately classify and route incoming tickets to the correct queue. 2. Automatically draft context-aware first responses using our proven saved replies and historical resolutions.
2.3 Core Constraint (Legal Compliance)
- Zero Automated Sending: Legal has mandated a strict human-in-the-loop requirement. The AI system must never autonomously send messages to customers.
- Financial Safety: The model is strictly prohibited from making automated commitments regarding refunds, credits, or billing adjustments without explicit human review and approval.
---
3. Goals and Success Metrics (KPIs)
| Metric | Current Baseline | Target (3 Months Post-Launch) |
|---|---|---|
| :--- | :--- | :--- |
| Average First Response Time (FRT) | 7 hours | < 2 hours |
| Routing Accuracy | ~75% (Manual/Rule-based)1 | > 95% across all four queues |
| Draft Acceptance Rate | N/A | > 70% (used as-is or with minor edits) |
| Agent Handle Time (AHT) | Baseline | -25% reduction in time spent crafting initial replies |
| Compliance / Unauthorized Sends | 0 | 0 violations (Hard safety guardrail) |
---
4. User Personas & Workflows
4.1 Support Agent
- Role: Resolves customer queries, reviews AI drafts, edits when necessary, approves responses, and handles complex escalations.
- Workflow: Opens the ticketing dashboard, sees incoming tickets pre-sorted into their designated queue with an AI-generated draft response attached. The agent reviews the draft, clicks "Approve & Send" or modifies the text, ensuring zero financial compliance risk.
4.2 Support Operations Lead
- Role: Manages queue health, tracks SLAs, audits agent performance, and maintains the library of saved replies and routing rules.
- Workflow: Monitors routing precision dashboards, updates macro/saved reply templates, and analyzes draft acceptance/rejection feedback loops to refine prompts.
4.3 Engineering & ML Leads
- Role: Build, deploy, scale, and monitor the inference pipeline, classification models, LLM orchestration, and UI integration.
---
5. Functional Requirements
5.1 Automated Ticket Routing (Triage Engine)
- Queue Classification: Upon ticket ingestion, the ML classification service must categorize incoming text into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
- Confidence Thresholds:
- Confidence $\ge$ 85%: Automatically assign the ticket to the predicted queue.
- Confidence < 85%: Route the ticket to a "Triage Review" sub-queue or default General queue with a visual tag flagging low classification confidence.
- Metadata Tagging: Tag tickets with sub-intent labels (e.g., `Billing: Refund Request`, `Technical: Bank Sync Failure`) to assist the response generation engine.
5.2 Response Draft Generation
- Context Window & Retrieval: The generation engine must leverage the incoming user query, customer account metadata (plan tier, active status), and relevant historical examples from the 2-year resolved ticket dataset (via RAG / vector search).
- Saved Reply Integration: The model must prioritize mapping inquiries to Ledgerly's 40 core saved replies where applicable, maintaining brand voice, accuracy, and tone consistency.
- Draft Status: Every generated response must be marked with a distinct internal status: `Draft - Pending Agent Review`. It must remain locked from customer view until explicit human sign-off.
5.3 Agent Workspace UI/UX Integration
- Side-by-Side Review: The agent UI must display the incoming customer message alongside the AI-generated draft in a clear, editable text box.
- Action Controls:
- [Approve & Send]: Immediately transmits the draft to the customer and logs the interaction.
- [Edit & Send]: Allows agents to modify text inline before sending. Edits must be logged for ML fine-tuning feedback loops.
- [Discard & Rewrite]: Clears the draft if the AI misunderstood the prompt.
- Visual Safety Badges: Clear UI warnings must appear on any ticket flagged as containing billing or refund keywords, reminding agents of compliance protocols.
---
6. Non-Functional & Safety Requirements
6.1 Legal & Financial Guardrails (Critical)
- No Auto-Dispatch: Zero API endpoints or automation rules are permitted to transition a draft state to `Sent` without a cryptographic token or database action originating from an authenticated human agent session.
- Refund Keyword Interceptor: A deterministic rule-based filter must scan all generated drafts for high-risk financial terminology (`refund`, `credit`, `waiver`, `chargeback`). If detected, the UI must highlight the section in amber and require a secondary manual checkbox confirmation from the agent before sending.
6.2 Performance & Latency
- Inference Latency: End-to-end classification and draft generation must complete within < 3.0 seconds of ticket ingestion via webhook, ensuring drafts are ready before an agent opens the ticket.
- Availability: 99.9% uptime SLA for the ML inference and routing service.
6.3 Data Privacy & Security
- PII Redaction: Customer tax IDs, banking credentials, and sensitive financial figures must be handled in compliance with GDPR, CCPA, and Ledgerly security policies. Training data pipelines must scrub extraneous PII prior to embedding generation.
---
7. Data, ML, & Training Strategy
7.1 Dataset Utilization
- Corpus: 2 years of historical resolved tickets (approx. 900,000+ interaction pairs across 42 agents).
- Data Splitting: 80% training/fine-tuning, 10% validation, 10% held-out test set reflecting recent seasonal billing spikes.
7.2 Model Architecture & Approach
- Classification: Fine-tuned encoder model (e.g., BERT/DeBERTa variant) for multi-class queue routing and intent tagging.
- Generation: Instruction-tuned LLM accessed via secure API or private VPC deployment, augmented with a vector database (RAG) containing the 40 approved saved replies and top-performing historical resolutions.
- Feedback Loop: Log all agent edits (diffs) and rejections to construct a weekly RLHF (Reinforcement Learning from Human Feedback) or supervised fine-tuning dataset to reduce hallucination rates.
---
8. High-Level Architecture & Data Flow
``` [Customer Email / Web Form] │ ▼ [Ingestion Service / Webhook] │ ├──► [ML Classification Engine] ──► Assigns Queue & Sub-intent │ └──► [RAG Vector DB (Saved Replies & History)] │ ▼ [LLM Draft Generation Service] ──► Generates Initial Response Text │ ▼ (Strictly Locked: Status = Pending Review) [Support Ticketing Database] │ ▼ [Agent UI Dashboard] ◄─── Human Agent Reviews, Edits, & Approves │ ▼ (Explicit Human Action) [Customer Dispatch API] ```
---
9. Phased Rollout Plan
- Phase 1: Shadow Mode (Weeks 1–4)
- Deploy classification and draft generation in the background.
- Log routing accuracy and draft quality without surfacing drafts to agents. Establish baseline ML accuracy metrics.
- Phase 2: Internal Pilot / Beta (Weeks 5–8)
- Roll out the UI feature to a pilot group of 5 senior Billing agents.
- Measure acceptance rates, time saved, and friction points. Refine prompt engineering and guardrails.
- Phase 3: General Availability Across All Queues (Weeks 9–12)2
- Roll out to all 42 agents across Billing, Technical, Account Access, and General queues.
- Establish continuous monitoring dashboards for Support Ops and Engineering leads.
---
10. Open Questions & Risks
- Edge Cases in Billing: How should the model handle complex multi-invoice dispute threads where historical context spans multiple months? (Mitigation: Surface the last 3 ticket summaries alongside the draft).
- Agent Adoption: Will agents trust the AI drafts, or will they rewrite them from scratch? (Mitigation: Emphasize time savings during training and incorporate agent feedback buttons).
What a PM had to fix
1Invented evidenceVerify or remove the claimStart again
Source: Volumes
9,000 tickets a week; 38% are billing.
What we’d changeRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.
2Test or gate too weakTighten the testSubstantial rework
What we’d changeAdvance each phase on quality gates (routing precision, draft acceptance, zero refund-commitment misses) rather than the calendar, with a human grading rubric and a rollback trigger.
Check by check
Got wrong · 6
- Uses the supplied evidence correctlyStates current routing accuracy of ~75% as fact without any support from the supplied evidence and does not label it as an assumption.
- Identifies material uncertaintyOpen questions are listed but are not tied to decision‑changing thresholds or explicit resolution plans; the PRD does not state what would halt or alter the rollout.
- Avoids unsupported claimsThe routing accuracy baseline of ~75% is presented as a fact without being marked as an assumption, and no evidence supports it.
- Defines evals and a launch barNo evaluation dataset criteria, metrics (e.g., draft quality), or specific launch threshold for progressing from shadow mode to pilot or GA are defined.
- Success settles trade-offsSuccess metrics include targets and a timeframe, but no explicit trade‑off rule (e.g., coverage vs. precision) is stated.
- Proposes tests that could failThe phased rollout lacks numeric thresholds, measurement windows, and actions tied to results; it does not define kill criteria.
Got right · 4
- Addresses the actual decisionThe PRD commits to a clear design for the AI feature, framed for the engineering, ML, and ops leads.
- Respects explicit constraintsThe document remains within the 1,000–1,500‑word range, addresses the required readers, and enforces legal constraints with human‑only dispatch and a refund‑keyword interceptor.
- Produces the required deliverableThe output is a complete PRD with functional requirements, architecture, and rollout plan that the named leads could start work from.
- Specifies AI failure behaviourLow-confidence classification routes to a triage queue, and a deterministic refund‑keyword interceptor triggers a secondary manual checkbox.
Claims the judge couldn’t find in the brief
- Current routing accuracy is ~75% (Manual/Rule-based).
Grades and run details
Decision model 50 · LLM judge 4 of 11 checks
Decision model checks
- failUses the supplied evidence correctly58%
- passAddresses the actual decision51%
- partialRespects explicit constraints40%
- partialIdentifies material uncertainty90%
- partialAvoids unsupported claims48%
- partialProduces the required deliverable46%
- partialSpecifies AI failure behaviour27%
- partialDefines evals and a launch bar83%
- partialSuccess settles trade-offs73%
- partialProposes tests that could fail69%
Artefacts
- link (link)
Run
- Run
- #1
- Time to output
- 14 s
- Submitted
- 24 Sept 2026