Needs more than a quick edit4 of 10 checks passed
A PM would need to remove or verify invented current-state facts and tighten the refund guardrail and length before using it. The reasoning is otherwise strong but not safe to ship as-is.
Output
Product Requirements Document (PRD)
Feature: AssistRoute (Automated Queue Routing & First-Response Drafting) Target Delivery: Q3 Status: Ready for Engineering Review Owners: Engineering Lead, ML Lead, Support Operations Lead
---
1. Problem Statement & Business Opportunity
Ledgerly processes 9,000 support tickets per week with a team of 42 agents across four queues: Billing (38%), Technical, Account Access, and General.
- Current FRT: 7.1 hours average first response time.
- Target FRT: < 2.0 hours.
- Manual Bottlenecks: Every incoming ticket requires manual reading, classification, and initial triage. Furthermore, approximately 60% of all initial inquiries map cleanly to one of ~40 canonical saved replies (macros), yet agents spend minutes manually searching, adapting, and populating them.
Solution Overview
AssistRoute is a two-part machine learning pipeline that acts on incoming tickets: 1. Predictive Routing: Automatically assigns incoming tickets to one of the four queues. 2. First-Response Drafting: Generates an editable, context-aware draft response pre-populated in the agent’s console before ticket open.
Core Invariant: Human-in-the-Loop (HITL). No draft is ever sent directly to a customer. Support agents retain 100% send authority.
---
2. Key Objectives & Metrics
| Metric Category | Target | Measurement Method |
|---|---|---|
| :--- | :--- | :--- |
| First Response Time (FRT) | $\le 2.0\text{ hours}$ (across all queues) | Timestamp delta: `ticket.createdat` to `firstagentmessage.sentat`. |
| Routing Accuracy | $\ge 93\%$ overall accuracy | Evaluated against tickets reassigned to a different queue within 24h. |
| Billing Routing Accuracy | $\ge 95\%$ precision/recall | Specific tracking on Billing due to volume (38%). |
| Draft Utilization Rate | $\ge 65\%$ of tickets | Proportion of first responses where the agent accepts the draft (as-is or edited). |
| Draft Edit Distance | $\le 30\%$ Levenshtein edit distance | Measures generation quality on accepted drafts. |
| Zero Refund Violation | 0 instances | Hard constraint: Zero drafts promising or confirming refunds reach customers without human authorization. |
---
3. Scope & Non-Goals
In Scope
- Asynchronous classification and routing pipeline for all newly created tickets via email and web form.
- Generation of a single personalized initial draft based on historical resolutions and the 40 standard macros.
- Guardrail pipeline strictly forbidding autonomous refund/credit commitments.
- Feedback loop telemetry (tracking agent accepts, edits, discards, and re-routes).
Out of Scope (Phase 1)
- Autonomous sending (auto-resolution without human click).
- Chat/live messaging support (limited strictly to asynchronous ticketing channels).
- Processing multi-turn responses (AssistRoute generates first responses only).
- Multi-language support (English only).
---
4. System Architecture & Workflow
``` [Customer Submits Ticket] │ ▼ [Event Ingestion: Webhook] ──▶ [PII Masking & Sanitization] │ ┌───────────────────────┴───────────────────────┐ ▼ ▼ [Routing Classifier] [Draft Generation Pipeline] │ │ Confidence $\ge$ Threshold? │ ├── Yes ──▶ Assign Queue 1. Macro Match / Retrieval └── No ──▶ Assign "General" (Flagged) 2. Prompt Compilation (LLM) │ 3. Deterministic Policy Guardrails │ │ └───────────────────────┬───────────────────────┘ ▼ [Agent Console: Pre-populated View] ├── Action A: Accept & Send ├── Action B: Edit & Send ├── Action C: Discard Draft └── Action D: Re-assign Queue ```
End-to-End Latency SLA
From `ticket.created` webhook ingestion to draft persistence in the ticketing database: $\le 5.0$ seconds (P95).
---
5. Functional Requirements
5.1 Routing Engine (ML Service)
- FR-1.1: The model must classify each ticket into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
- FR-1.2: Confidence Scoring:
- If `confidence >= 0.85`: Automatically set the ticket's `queue_id` attribute.
- If `confidence < 0.85`: Set `queueid` to `General` and append the tag `needstriage`.
- FR-1.3: The Routing Engine must evaluate and route the ticket before generating the draft, as draft prompts require queue-specific system contexts.
5.2 Retrieval & Generation Pipeline
- FR-2.1 (Context Hydration): The generation service must pull:
- The ticket subject and body.
- The authenticated user's metadata: Plan tier (`Solo`, `Growth`, `Enterprise`), account age, and active add-on modules.
- The closest semantic match from Ledgerly's 40 standard macros.
- FR-2.2 (Draft Generation): The system must generate a friendly, concise first response that adheres to Ledgerly brand voice, incorporates the customer's name, and directly addresses the primary issue using the retrieved macro logic.
- FR-2.3 (Confidence Gating): If the model's semantic similarity score against approved knowledge bases or historical solutions falls below `0.70`, no draft shall be rendered. The UI must show: "Draft unavailable: Low context confidence."
5.3 Legal & Compliance Guardrail (Refund Protection)
* FR-3.1 (Deterministic Regex & Semantic Filter): Every draft must pass through a two-stage financial commitment filter prior to saving: 1. Lexical Check: Negative keyword blocking: `refund`, `credit`, `reimburse`, `waive fee`, `reverse charge`, `money back`, `comp`. 2. Semantic Entailment Check: Verification that the text does not commit Ledgerly to monetary compensation or liability. * FR-3.2 (Violation Action): If a draft contains a prohibited financial commitment: * The drafting engine must automatically strip the phrase and substitute safe boilerplate: "I have escalated your billing inquiry to our finance specialists for immediate review." * The draft must display a visible amber badge in the Agent UI: "Financial Review Required — No Commitments Permitted."
5.4 Support Agent Console UX
- FR-4.1 (Draft Presentation): The draft must render inside the native response editor as an uncommitted state (distinct background styling or watermarked).
- FR-4.2 (Interaction Actions):
- Tab / Single-Click Acceptance: Agent hits `Tab` or clicks `Use Draft` to copy the generated text into the active editor.
- Direct Editing: Editing the draft directly must be seamless without extra confirmation modals.
- Clear / Reject: Single-click action to clear the generated text and open a blank response field.
- FR-4.3 (Metadata Panel): A sidebar panel must show:
- Queue classification and confidence score.
- Which of the 40 macros informed the draft.
- A one-click queue reassignment dropdown.
``` +-------------------------------------------------------------------------+ | Ticket #10492: "Charged twice for Payroll add-on" | | Sender: Sarah Jenkins (Growth Tier) | Queue: [Billing (96%)] [Re-route] | +-------------------------------------------------------------------------+ | [ AI Drafted Response ] [Macro Source: Billing_Dup]| | | | Hi Sarah, | | | | Thanks for reaching out. I see that your account was debited twice | | for the Payroll add-on on October 1st. | | | | I have escalated this to our payment operations team to review the | | duplicate transaction details. We will update you here as soon as this | | review is complete. | | | | [ Accept Draft (Tab) ] [ Edit Draft ] [ Reject ] | +-------------------------------------------------------------------------+ | [ Agent Response Box ] | +-------------------------------------------------------------------------+ ```
---
6. ML & Data Engineering Specifications
6.1 Data Preparation (2-Year Historical Corpus)
1. Sanitization: Strip all historical PII (tax identifiers, SSNs, credit card numbers, passwords) using Microsoft Presidio or an equivalent NER pipeline before training/indexing. 2. Filtering: * Drop all historical tickets that required more than 4 re-routes (noisy labels). * Filter out tickets closed with negative customer satisfaction (CSAT $\le 2$). * Exclude responses superseded by outdated accounting rules or old pricing tiers (Ops team to define date cutoffs). 3. Macro Ground-Truth Alignment: Map the 40 canonical macros against historical agent responses to serve as gold-standard reference pairs.
6.2 Model Specifications
- Routing Classifier:
- Architecture: Fine-tuned lightweight encoder (e.g., `modern-bert-base` or `RoBERTa-base`) or an optimized classification endpoint.
- Input: Ticket Subject + Body.
- Output: Softmax distribution over `[Billing, Technical, Account Access, General]`.
- Latency Target: $< 200\text{ ms}$.
- Generative Drafting Model:
- Architecture: Hosted LLM (e.g., Claude 3.5 Sonnet or GPT-4o-mini) via secure enterprise API with zero-data-retention agreements.
- Prompt Design: System prompt containing Ledgerly tone guidelines, user context JSON, the selected macro instructions, and strict instructions forbidding financial promises.
- Temperature: `0.1` (low variability, high determinism).
6.3 Telemetry & Event Logging
The frontend and backend must emit the following events to the analytics lakehouse: * `ticketrouted`: `{ ticketid, predictedqueue, confidence, autoassigned: bool }` * `ticketrerouted`: `{ ticketid, oldqueue, newqueue, agentid }` * `draftgenerated`: `{ ticketid, macroid, promptversion, modelid, generationtimems }` * `draftactioned`: * `action`: `ACCEPTEDASIS` | `EDITED` | `REJECTED` * `originaldraft`: string * `finalsenttext`: string * `levenshteindistance`: int * `agentid`: string
---
7. Safety, Compliance, and Security
- Refund Protection Hard Filter: As specified in Section 5.3, any draft containing automated refund promises is blocked from the UI by an air-gapped deterministic validation check that runs after LLM generation and cannot be bypassed by prompt engineering.
- Zero Training on Customer Data: LLM vendors must have explicit Data Processing Agreements (DPAs) confirming customer ticket content is not used to train foundation models.
- Data Storage & SOC 2: Masked inputs and outputs stored for model evaluation must adhere to standard data retention policies (auto-purged after 90 days).
---
8. Rollout Plan & Milestones
``` Phase 0: Offline Validation & Guardrail Auditing (Weeks 1-3) ├── Train routing model on historical data; evaluate against held-out test set. ├── Run 10,000 synthetic adversarial prompts targeting the refund guardrail. └── Success Gate: Routing accuracy > 93%, 0 adversarial refund breaches.
Phase 1: Shadow Mode (Weeks 4-5) ├── Pipeline processes live tickets; outputs are logged to Databricks/Snowflake. ├── No UI visible to agents. └── Evaluate: Model inference latency, real-time routing accuracy vs. manual triage.
Phase 2: Single-Queue Canary Pilot (Weeks 6-7) ├── Deploy exclusively to the Billing Queue (10 designated agents). ├── Daily sync with Support Operations leads to review draft quality. └── Success Gate: Draft utilization > 50%, no severe agent blockers.
Phase 3: General Availability (GA) & Iteration (Weeks 8+) ├── Enable routing and drafting across all 42 agents and 4 queues. ├── Implement weekly retraining cadence for classifier based on override logs. └── Track progress toward < 2.0 hour FRT target. ```
---
9. Risks & Mitigations
| Risk | Impact | Likelihood | Mitigation Strategy |
|---|---|---|---|
| :--- | :--- | :--- | :--- |
| Agent over-reliance (rubber-stamping bad drafts) | High | Medium | Random 2% QA audits on accepted drafts by Support Operations leads. Prominent UI reminders that agents are solely responsible for sent text. |
| Hallucinated accounting/tax guidance | Critical | Low | Restrict generation strictly to standard macros and product knowledge base; temperature set to 0.1; fallback to "Draft unavailable" when query context is ambiguous. |
| Model drift due to product updates | Medium | High | Weekly regression testing on macro sets. Support Ops manages a macro repository that directly updates retrieval indexes without code deploys. |
| Routing loop (agent re-routes ticket repeatedly) | Low | Low | Hard ceiling: If a ticket is manually re-routed twice, AssistRoute disables auto-routing for that ticket and locks it to manual handling. |
Claims the judge couldn’t find in the brief
- Current average first response time is 7.1 hours.
- Every incoming ticket requires manual reading, classification, and initial triage.
- Agents spend minutes manually searching, adapting, and populating saved replies.
- New tickets arrive via email and web form.
- Authenticated user metadata includes plan tier (Solo, Growth, Enterprise), account age, and active add-on modules.
- Historical tickets contain PII such as tax identifiers, SSNs, credit card numbers, and passwords.
- Historical tickets include CSAT scores and re-route counts.
- The company uses Databricks/Snowflake and has SOC 2 controls.
Grades and run details
Decision model 65 · LLM judge 4 of 11 checks
Decision model checks
- failUses the supplied evidence correctly59%
- passAddresses the actual decision87%
- partialRespects explicit constraints26%
- partialIdentifies material uncertainty34%
- partialAvoids unsupported claims30%
- passProduces the required deliverable55%
- passSpecifies AI failure behaviour98%
- passDefines evals and a launch bar92%
- partialSuccess settles trade-offs75%
- partialProposes tests that could fail69%
Run
- Run
- #1
- API response time
- 33 s
- Submitted
- 1 Oct 2026