Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given3 items: Scenario, Volumes, Risk
ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer doesThe answer key the graders mark against

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score at 40)
  • Allows automated refund commitments
Case

Staff-level · v1.11 · synthetic · AI product, support

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 10 checks passed

A PM would need to remove or verify invented current-state facts and tighten the refund guardrail and length before using it. The reasoning is otherwise strong but not safe to ship as-is.

Output

Product Requirements Document (PRD)

Feature: AssistRoute (Automated Queue Routing & First-Response Drafting) Target Delivery: Q3 Status: Ready for Engineering Review Owners: Engineering Lead, ML Lead, Support Operations Lead

---

1. Problem Statement & Business Opportunity

Ledgerly processes 9,000 support tickets per week with a team of 42 agents across four queues: Billing (38%), Technical, Account Access, and General.

  • Current FRT: 7.1 hours average first response time.
  • Target FRT: < 2.0 hours.
  • Manual Bottlenecks: Every incoming ticket requires manual reading, classification, and initial triage. Furthermore, approximately 60% of all initial inquiries map cleanly to one of ~40 canonical saved replies (macros), yet agents spend minutes manually searching, adapting, and populating them.

Solution Overview

AssistRoute is a two-part machine learning pipeline that acts on incoming tickets: 1. Predictive Routing: Automatically assigns incoming tickets to one of the four queues. 2. First-Response Drafting: Generates an editable, context-aware draft response pre-populated in the agent’s console before ticket open.

Core Invariant: Human-in-the-Loop (HITL). No draft is ever sent directly to a customer. Support agents retain 100% send authority.

---

2. Key Objectives & Metrics

Metric CategoryTargetMeasurement Method
:---:---:---
First Response Time (FRT)$\le 2.0\text{ hours}$ (across all queues)Timestamp delta: `ticket.createdat` to `firstagentmessage.sentat`.
Routing Accuracy$\ge 93\%$ overall accuracyEvaluated against tickets reassigned to a different queue within 24h.
Billing Routing Accuracy$\ge 95\%$ precision/recallSpecific tracking on Billing due to volume (38%).
Draft Utilization Rate$\ge 65\%$ of ticketsProportion of first responses where the agent accepts the draft (as-is or edited).
Draft Edit Distance$\le 30\%$ Levenshtein edit distanceMeasures generation quality on accepted drafts.
Zero Refund Violation0 instancesHard constraint: Zero drafts promising or confirming refunds reach customers without human authorization.

---

3. Scope & Non-Goals

In Scope

  • Asynchronous classification and routing pipeline for all newly created tickets via email and web form.
  • Generation of a single personalized initial draft based on historical resolutions and the 40 standard macros.
  • Guardrail pipeline strictly forbidding autonomous refund/credit commitments.
  • Feedback loop telemetry (tracking agent accepts, edits, discards, and re-routes).

Out of Scope (Phase 1)

  • Autonomous sending (auto-resolution without human click).
  • Chat/live messaging support (limited strictly to asynchronous ticketing channels).
  • Processing multi-turn responses (AssistRoute generates first responses only).
  • Multi-language support (English only).

---

4. System Architecture & Workflow

``` [Customer Submits Ticket] │ ▼ [Event Ingestion: Webhook] ──▶ [PII Masking & Sanitization] │ ┌───────────────────────┴───────────────────────┐ ▼ ▼ [Routing Classifier] [Draft Generation Pipeline] │ │ Confidence $\ge$ Threshold? │ ├── Yes ──▶ Assign Queue 1. Macro Match / Retrieval └── No ──▶ Assign "General" (Flagged) 2. Prompt Compilation (LLM) │ 3. Deterministic Policy Guardrails │ │ └───────────────────────┬───────────────────────┘ ▼ [Agent Console: Pre-populated View] ├── Action A: Accept & Send ├── Action B: Edit & Send ├── Action C: Discard Draft └── Action D: Re-assign Queue ```

End-to-End Latency SLA

From `ticket.created` webhook ingestion to draft persistence in the ticketing database: $\le 5.0$ seconds (P95).

---

5. Functional Requirements

5.1 Routing Engine (ML Service)

  • FR-1.1: The model must classify each ticket into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • FR-1.2: Confidence Scoring:
  • If `confidence >= 0.85`: Automatically set the ticket's `queue_id` attribute.
  • If `confidence < 0.85`: Set `queueid` to `General` and append the tag `needstriage`.
  • FR-1.3: The Routing Engine must evaluate and route the ticket before generating the draft, as draft prompts require queue-specific system contexts.

5.2 Retrieval & Generation Pipeline

  • FR-2.1 (Context Hydration): The generation service must pull:
  • The ticket subject and body.
  • The authenticated user's metadata: Plan tier (`Solo`, `Growth`, `Enterprise`), account age, and active add-on modules.
  • The closest semantic match from Ledgerly's 40 standard macros.
  • FR-2.2 (Draft Generation): The system must generate a friendly, concise first response that adheres to Ledgerly brand voice, incorporates the customer's name, and directly addresses the primary issue using the retrieved macro logic.
  • FR-2.3 (Confidence Gating): If the model's semantic similarity score against approved knowledge bases or historical solutions falls below `0.70`, no draft shall be rendered. The UI must show: "Draft unavailable: Low context confidence."

5.3 Legal & Compliance Guardrail (Refund Protection)

* FR-3.1 (Deterministic Regex & Semantic Filter): Every draft must pass through a two-stage financial commitment filter prior to saving: 1. Lexical Check: Negative keyword blocking: `refund`, `credit`, `reimburse`, `waive fee`, `reverse charge`, `money back`, `comp`. 2. Semantic Entailment Check: Verification that the text does not commit Ledgerly to monetary compensation or liability. * FR-3.2 (Violation Action): If a draft contains a prohibited financial commitment: * The drafting engine must automatically strip the phrase and substitute safe boilerplate: "I have escalated your billing inquiry to our finance specialists for immediate review." * The draft must display a visible amber badge in the Agent UI: "Financial Review Required — No Commitments Permitted."

5.4 Support Agent Console UX

  • FR-4.1 (Draft Presentation): The draft must render inside the native response editor as an uncommitted state (distinct background styling or watermarked).
  • FR-4.2 (Interaction Actions):
  • Tab / Single-Click Acceptance: Agent hits `Tab` or clicks `Use Draft` to copy the generated text into the active editor.
  • Direct Editing: Editing the draft directly must be seamless without extra confirmation modals.
  • Clear / Reject: Single-click action to clear the generated text and open a blank response field.
  • FR-4.3 (Metadata Panel): A sidebar panel must show:
  • Queue classification and confidence score.
  • Which of the 40 macros informed the draft.
  • A one-click queue reassignment dropdown.

``` +-------------------------------------------------------------------------+ | Ticket #10492: "Charged twice for Payroll add-on" | | Sender: Sarah Jenkins (Growth Tier) | Queue: [Billing (96%)] [Re-route] | +-------------------------------------------------------------------------+ | [ AI Drafted Response ] [Macro Source: Billing_Dup]| | | | Hi Sarah, | | | | Thanks for reaching out. I see that your account was debited twice | | for the Payroll add-on on October 1st. | | | | I have escalated this to our payment operations team to review the | | duplicate transaction details. We will update you here as soon as this | | review is complete. | | | | [ Accept Draft (Tab) ] [ Edit Draft ] [ Reject ] | +-------------------------------------------------------------------------+ | [ Agent Response Box ] | +-------------------------------------------------------------------------+ ```

---

6. ML & Data Engineering Specifications

6.1 Data Preparation (2-Year Historical Corpus)

1. Sanitization: Strip all historical PII (tax identifiers, SSNs, credit card numbers, passwords) using Microsoft Presidio or an equivalent NER pipeline before training/indexing. 2. Filtering: * Drop all historical tickets that required more than 4 re-routes (noisy labels). * Filter out tickets closed with negative customer satisfaction (CSAT $\le 2$). * Exclude responses superseded by outdated accounting rules or old pricing tiers (Ops team to define date cutoffs). 3. Macro Ground-Truth Alignment: Map the 40 canonical macros against historical agent responses to serve as gold-standard reference pairs.

6.2 Model Specifications

  • Routing Classifier:
  • Architecture: Fine-tuned lightweight encoder (e.g., `modern-bert-base` or `RoBERTa-base`) or an optimized classification endpoint.
  • Input: Ticket Subject + Body.
  • Output: Softmax distribution over `[Billing, Technical, Account Access, General]`.
  • Latency Target: $< 200\text{ ms}$.
  • Generative Drafting Model:
  • Architecture: Hosted LLM (e.g., Claude 3.5 Sonnet or GPT-4o-mini) via secure enterprise API with zero-data-retention agreements.
  • Prompt Design: System prompt containing Ledgerly tone guidelines, user context JSON, the selected macro instructions, and strict instructions forbidding financial promises.
  • Temperature: `0.1` (low variability, high determinism).

6.3 Telemetry & Event Logging

The frontend and backend must emit the following events to the analytics lakehouse: * `ticketrouted`: `{ ticketid, predictedqueue, confidence, autoassigned: bool }` * `ticketrerouted`: `{ ticketid, oldqueue, newqueue, agentid }` * `draftgenerated`: `{ ticketid, macroid, promptversion, modelid, generationtimems }` * `draftactioned`: * `action`: `ACCEPTEDASIS` | `EDITED` | `REJECTED` * `originaldraft`: string * `finalsenttext`: string * `levenshteindistance`: int * `agentid`: string

---

7. Safety, Compliance, and Security

  1. Refund Protection Hard Filter: As specified in Section 5.3, any draft containing automated refund promises is blocked from the UI by an air-gapped deterministic validation check that runs after LLM generation and cannot be bypassed by prompt engineering.
  2. Zero Training on Customer Data: LLM vendors must have explicit Data Processing Agreements (DPAs) confirming customer ticket content is not used to train foundation models.
  3. Data Storage & SOC 2: Masked inputs and outputs stored for model evaluation must adhere to standard data retention policies (auto-purged after 90 days).

---

8. Rollout Plan & Milestones

``` Phase 0: Offline Validation & Guardrail Auditing (Weeks 1-3) ├── Train routing model on historical data; evaluate against held-out test set. ├── Run 10,000 synthetic adversarial prompts targeting the refund guardrail. └── Success Gate: Routing accuracy > 93%, 0 adversarial refund breaches.

Phase 1: Shadow Mode (Weeks 4-5) ├── Pipeline processes live tickets; outputs are logged to Databricks/Snowflake. ├── No UI visible to agents. └── Evaluate: Model inference latency, real-time routing accuracy vs. manual triage.

Phase 2: Single-Queue Canary Pilot (Weeks 6-7) ├── Deploy exclusively to the Billing Queue (10 designated agents). ├── Daily sync with Support Operations leads to review draft quality. └── Success Gate: Draft utilization > 50%, no severe agent blockers.

Phase 3: General Availability (GA) & Iteration (Weeks 8+) ├── Enable routing and drafting across all 42 agents and 4 queues. ├── Implement weekly retraining cadence for classifier based on override logs. └── Track progress toward < 2.0 hour FRT target. ```

---

9. Risks & Mitigations

RiskImpactLikelihoodMitigation Strategy
:---:---:---:---
Agent over-reliance (rubber-stamping bad drafts)HighMediumRandom 2% QA audits on accepted drafts by Support Operations leads. Prominent UI reminders that agents are solely responsible for sent text.
Hallucinated accounting/tax guidanceCriticalLowRestrict generation strictly to standard macros and product knowledge base; temperature set to 0.1; fallback to "Draft unavailable" when query context is ambiguous.
Model drift due to product updatesMediumHighWeekly regression testing on macro sets. Support Ops manages a macro repository that directly updates retrieval indexes without code deploys.
Routing loop (agent re-routes ticket repeatedly)LowLowHard ceiling: If a ticket is manually re-routed twice, AssistRoute disables auto-routing for that ticket and locks it to manual handling.

Claims the judge couldn’t find in the brief

  • Current average first response time is 7.1 hours.
  • Every incoming ticket requires manual reading, classification, and initial triage.
  • Agents spend minutes manually searching, adapting, and populating saved replies.
  • New tickets arrive via email and web form.
  • Authenticated user metadata includes plan tier (Solo, Growth, Enterprise), account age, and active add-on modules.
  • Historical tickets contain PII such as tax identifiers, SSNs, credit card numbers, and passwords.
  • Historical tickets include CSAT scores and re-route counts.
  • The company uses Databricks/Snowflake and has SOC 2 controls.

Grades and run details

Decision model 65 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly59%
  • passAddresses the actual decision87%
  • partialRespects explicit constraints26%
  • partialIdentifies material uncertainty34%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable55%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar92%
  • partialSuccess settles trade-offs75%
  • partialProposes tests that could fail69%
Run
Run
#1
API response time
33 s
Submitted
1 Oct 2026

Needs more than a quick edit7 of 10 checks passed

It is largely usable but invents current systems not in the brief, so a PM would need to verify the evidence before acting. It also lacks an explicit trade-off rule and complete test windows/actions, requiring more than a quick edit.

Output

PRD: AI-Assisted First Response and Ticket Routing

Product: Ledgerly Support Assist Audience: Engineering, ML and Support Operations leads Status: Proposed MVP; thresholds below are launch targets, not measured results Accountable owner: Support Operations lead, with Engineering and ML leads owning delivery and model quality

1. Problem and objective

Ledgerly’s 42 support agents handle 9,000 tickets per week across Billing, Technical, Account access and General. Billing represents 38% of volume—approximately 3,420 tickets weekly. Average first-response time is seven hours, against a target of under two hours. Approximately 60% of tickets receive one of roughly 40 saved replies.

The feature will classify new tickets into the appropriate queue and prepare a grounded first-response draft for an agent to review. It should reduce queue-selection work and repetitive writing without delegating customer communication or financial commitments to the model.

Every outbound response requires explicit agent approval. Legal prohibits automated sending of refund commitments. The MVP will have no automated-send path and will not generate refund promises; agents must author any commitment using the existing authorized refund process.

2. Goals and success measures

Primary goals:

  • Reduce average time to the first substantive human-approved response to under two hours during a staffed pilot.
  • Reduce median active agent time spent preparing first responses by at least 30%.
  • Route tickets accurately while avoiding hidden misroutes.
  • Preserve response correctness, account security and customer satisfaction.

Measure first-response time from ticket creation to the first substantive response sent by an agent. Automated receipts do not count. Report calendar-hour and staffed-hour results separately, plus median and p90, to prevent averages hiding long waits.

Additional proposed launch and pilot thresholds:

MeasureTarget
Queue-classification accuracy≥95% on the adjudicated holdout
Account access routing recall≥98%
High-confidence automatic-routing precision≥98%
Drafts usable with no substantive correction≥80% in blinded review
Unsupported financial or security commitmentsZero observed in launch evaluation
Draft availability latencyp95 ≤15 seconds after ingestion
Pilot qualityNo material deterioration in QA scores or customer satisfaction

Zero observed errors is a release gate, not proof of zero production risk. ML must report sample sizes and confidence intervals. Support Operations must confirm staffing coverage: drafting improvements alone cannot guarantee the response-time target.

3. Scope

MVP includes:

  • Classification into Billing, Technical, Account access or General.
  • Automatic routing only above validated, queue-specific confidence thresholds.
  • Retrieval of approved saved replies and current support documentation.
  • A first-response draft with internal source references, risk flags and suggested clarifying questions.
  • Agent controls to edit, discard, regenerate, change queue and approve/send.
  • Audit logs, monitoring and immediate disable controls.

Excluded: automatic sending, subsequent-turn assistance, ticket resolution, refunds, account changes, security verification decisions and customer-facing AI chat.

Ticket channels, languages and attachment formats are not specified. Support Operations will inventory them before implementation. MVP supports only validated text channels and languages; unsupported inputs receive normal human triage without a generated draft.

4. User workflow and functional requirements

  1. Ingest: On creation of a new eligible ticket, capture its text, subject and approved metadata. Use only authorized customer/account context available to the assigned support role.
  2. Assess: Predict a queue, calibrated confidence and risk flags such as refund request, account compromise or insufficient information.
  3. Route: Assign high-confidence tickets to the predicted queue. Send uncertain cases to the existing manual-triage destination, provisionally General. Surface uncertainty prominently; do not treat General as a confident classification.
  4. Draft: Retrieve relevant approved material and produce a concise response addressing the request. Prefer adapting an applicable saved reply over open-ended generation.
  5. Review: Show the draft, suggested queue, risk flags and source links in the agent workspace. Clearly label the text “AI draft—not sent.”
  6. Approve/send: The existing send action requires an explicit authenticated agent action. No background job, timeout or model output may trigger sending.
  7. Learn: Record queue corrections, edits, rejection reasons and quality reviews for controlled evaluation and future retraining.

Drafts must not invent transactions, troubleshooting outcomes, refund eligibility or account status. Where evidence is missing, ask for necessary information or acknowledge that an agent must investigate. Avoid requesting passwords, full payment-card details or other unnecessary sensitive data.

Refund-related drafts may acknowledge the request and explain approved next steps, but must not promise an amount, eligibility or processing date. Flag these tickets for Billing review. Agents may add commitments only after authorized verification.

A draft becomes stale if relevant ticket content or account context changes. Disable approval until it is refreshed or explicitly reviewed against the updated context. Manual queue changes take precedence over later model results.

5. Data and ML approach

Two years of resolved tickets and agent replies are available. Historical replies are examples, not authoritative policy: outdated guidance and unauthorized commitments must not be reproduced.

ML and Support Operations will:

  • Define queue labels and rules for mixed-intent tickets. Account-security concerns take precedence over routine billing questions; Support Operations must approve the complete precedence matrix.
  • Audit final queue labels and sample ambiguous cases for expert adjudication.
  • Remove duplicates, signatures and irrelevant quoted history; redact unnecessary personal and payment information.
  • Preserve ticket/thread/customer grouping across splits to reduce leakage.
  • Use a chronological training, validation and held-out test split, with the newest period reserved for testing.
  • Evaluate only information available when the ticket arrived. Later replies and resolutions may provide labels but must never become runtime inputs.

Benchmark saved-reply retrieval and a conventional classifier before introducing more complex models. Select the simplest approach meeting quality, latency and operational requirements.

Generation will use current, versioned approved content. Historical replies may support offline training subject to privacy and quality approval, but must not serve as an unrestricted runtime knowledge source. Conflicting or missing sources trigger a clarification or human-investigation draft rather than a guessed answer.

Customer text and retrieved content are untrusted inputs. Instructions embedded in tickets must not change system rules, authorize actions or bypass review.

6. Engineering design and controls

Implement an asynchronous pipeline behind feature flags:

Ticket event → eligibility check → classifier → routing decision → retrieval → generation → policy validation → draft storage → agent UI.

Persist ticket ID/version, model and prompt versions, retrieved document versions, queue scores, draft state, risk flags and agent actions. Use idempotency keys to prevent duplicate processing and concurrency controls to avoid overwriting agent work.

The AI service must have no credential or permission to send messages, issue refunds or modify accounts. Routing permissions must be limited to approved queues. Outbound messages remain controlled by the existing authenticated support application.

Post-generation checks will block prohibited promises and unsupported sensitive claims. Failed checks suppress the draft and display a reason; they must not merely append a disclaimer.

Apply role-based access, encryption and existing retention rules. Any external model provider requires Security and Legal approval, including contractual restrictions on retention and training use. Log identifiers and operational metadata where possible, not unrestricted ticket bodies.

On timeout, provider outage or validation failure, preserve normal ticket handling. Keep the ticket visible, apply manual triage where necessary and show “Draft unavailable.” Never delay ticket intake while waiting for AI.

7. Evaluation and release gates

Build an adjudicated test set covering all queues, common saved-reply cases, mixed intent, sparse descriptions, refund disputes, account compromise, outdated-policy examples and prompt-injection attempts. Report overall and per-queue performance; oversampled risk cases must also be reported separately from production-weighted results.

Two support reviewers will score drafts for correctness, relevance, completeness, tone and policy compliance, with disagreements adjudicated. Distinguish cosmetic edits from substantive corrections. Track source support and appropriate abstention, not acceptance rate alone.

Required acceptance tests include:

  • A refund-request draft contains no commitment.
  • No model or pipeline component can invoke sending.
  • Unsupported facts cause omission or escalation.
  • Ticket updates invalidate stale drafts.
  • Human rerouting is not overwritten.
  • Provider failure leaves tickets accessible.
  • Duplicate events create neither duplicate drafts nor duplicate sends.
  • Agent permissions also constrain retrieved context.

Production expansion requires meeting quality thresholds and passing all safety-critical tests. Any unauthorized financial promise in evaluation blocks launch pending remediation.

8. Rollout, ownership and operations

Indicative eight-week plan:

  • Weeks 1–2: Support Operations approves taxonomy, saved replies, triage rules and baseline measurements; ML audits data; Engineering completes integration and security design.
  • Weeks 3–4: Build the pipeline, UI, evaluation harness and operational controls.
  • Weeks 5–6: Run offline evaluation and shadow mode without changing queues or showing drafts.
  • Weeks 7–8: Pilot with selected agents, then expand by queue only after gates pass.

Use a randomized holdout where practical, balanced by queue and shift, to compare handling time, response time, corrections and customer outcomes.

Engineering owns availability, access controls and rollback. ML owns evaluation, calibration and drift monitoring. Support Operations owns policy content, QA, agent training and staffing. Legal approves refund wording and provider terms.

Provide separate kill switches for routing and drafting. Suspected unauthorized commitments, data exposure or account-security mishandling trigger immediate suspension of the affected capability and incident review. Review corrections and rejected drafts weekly; retraining requires fresh holdout evaluation, not automatic ingestion of agent edits.

Claims the judge couldn’t find in the brief

  • The existing send action requires an explicit authenticated agent action.
  • There is an existing manual-triage destination, provisionally General.
  • There is an existing authorized refund process agents must use.

Grades and run details

Decision model 85 · LLM judge 7 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly7%
  • passAddresses the actual decision90%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty40%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable68%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • partialSuccess settles trade-offs27%
  • partialProposes tests that could fail70%
Run
Run
#1
API response time
49 s
Submitted
29 Sept 2026

Needs more than a quick edit8 of 10 checks passed

Unsupported claims about current workflow and data content would require re-checking evidence, and the data section's assumptions need validation before engineering can act.

Output

PRD: AI Draft & Route for Support Tickets

Sep 24, 2026 · @Martin Slaney

1. Summary and problem

Ledgerly's support team (42 agents, four queues: Billing, Technical, Account access, General) handles \~9,000 tickets a week. Average first response time (FRT) is 7 hours against a target of under 2.

Most of that delay is not writing time. It is triage and lookup1: \~60% of tickets are answered with one of \~40 saved replies, and 38% are billing. Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details.

This feature does two things when a ticket arrives:

  1. Routes it to the correct queue, with a confidence score.
  2. Drafts a first response, grounded in the saved-reply library, the customer's account data and similar past resolutions, for an agent to review.

Nothing is sent without an agent approving it. On top of that, any draft containing a refund commitment is hard-gated per Legal (§4.3).

2. Goals, non-goals and success metrics

Goals

  • Bring average FRT under 2 hours without adding headcount.
  • Cut misrouted tickets and the reassignment delay they cause.
  • Reduce agent handling time on saved-reply tickets, so agents spend more time on the \~40% that need real investigation.

Non-goals (v1)

  • Auto-sending any response, of any type.
  • Drafting follow-ups beyond the first response.
  • Taking actions in Ledgerly (issuing refunds, resetting passwords, changing plans). The model drafts text; agents act.
  • Customer-facing chatbot or deflection.

Success metrics (measured per queue, against a 4-week pre-launch baseline)

MetricTargetGuardrail
Average FRT< 2hP90 FRT must not rise
Routing accuracy (final queue = predicted queue)≥ 92% overall, ≥ 90% per queueAccount access recall ≥ 95%
Draft acceptance (sent with light or no edits)≥ 50% of drafted tickets—
Agent handling time on drafted tickets−30%—
CSAT on drafted ticketsNo drop vs baseline (±1pt)Reopen rate not up >1pt
Refund commitments sent without refund-approval step0Hard requirement

"Light edit" is defined as a normalised edit distance below 0.2 between draft and sent text. ML owns the metric definition; Support Ops signs it off before pilot.

3. Users and core workflow

Users: support agents (reviewers of every draft); queue leads (monitor routing, handle overrides); Support Ops (owns saved replies, policies and the refund-approval rota).

Flow for a new ticket

  1. Ticket created (email or in-app form).
  2. Router predicts queue + confidence within 30s. High confidence → assigned to that queue. Low confidence → Triage view for a lead to assign in one click.
  3. Drafter produces a first response and attaches it to the ticket as an internal draft, with: the saved reply(s) it drew on, account facts it inserted, and any flags (refund, low confidence, missing data).
  4. Agent opens the ticket, reviews the draft, and chooses Send, Edit & send, or Discard (with a reason code).
  5. If the draft is refund-flagged, Send is replaced by Request refund approval (§4.3).
  6. Final queue, sent text and action are logged as training and evaluation signal.

4. Functional requirements

4.1 Routing

  • R1. Classify every new ticket into Billing, Technical, Account access or General, with a calibrated confidence score.
  • R2. Auto-assign when confidence ≥ threshold (set per queue from eval, starting target: ≥ 95% precision at that threshold). Below threshold → Triage view.
  • R3. Account access is the costliest miss (locked-out customers).3 Tune for recall on this class; a ticket with any access signal and ambiguous classification goes to Account access, not General.
  • R4. Agents can reassign in one click; every reassignment is logged with the original prediction.
  • R5. Support Ops can switch routing to suggest-only per queue without a deploy.

4.2 Drafting

  • D1. Generate a draft for every routed ticket in English within 60s of creation. Other languages: no draft in v1, flag only.
  • D2. Retrieval first: identify the best-matching saved reply (or "none"). When one matches, the draft is that reply personalised with ticket and account context, not free text. When none matches, draft from similar resolved tickets and help-centre articles, and label it Free-form draft.
  • D3. Account facts (plan, billing dates, invoice amounts, last payment status) come only from read-only lookups against Ledgerly's billing/account APIs, never from model memory. Any fact the model could not verify is left as a visible `[placeholder]` that blocks sending until filled.
  • D4. Show sources inline: saved-reply ID, linked tickets/articles, API fields used.
  • D5. Never promise timelines, credits, discounts, policy exceptions or refunds unless the saved reply itself contains them2 (refunds additionally gated, §4.3).
  • D6. Skip drafting (flag only) for: legal threats, suspected fraud, data deletion/GDPR requests, security incidents, and abusive or distressed customers. Keyword + classifier; list owned by Support Ops.

4.3 Refund guardrail (Legal requirement)

Legal requires no automated sending of refund commitments. Because every draft already needs agent approval, we go further so a refund promise cannot slip through in a routine approve click:

  • G1. A dedicated refund-commitment detector runs on every draft and on the final edited text at send time (agents may add refund language themselves). Tuned for recall ≥ 99% on a Legal-reviewed test set; false positives are acceptable.
  • G2. If triggered, the ticket cannot be sent via one-click approval. The agent must confirm the refund is authorised under current policy (Support Ops defines who can authorise which amounts) via a separate confirmation step that is logged.
  • G3. The drafter must never generate a refund commitment from free-form reasoning; refund language may only come from approved refund saved replies.
  • G4. No bulk-approve action exists anywhere in the product.
  • G5. Legal reviews the detector test set and the confirmation UX before pilot, and receives a monthly log of refund-flagged sends.

Open for Legal: does "automated sending" cover a one-click approve by an agent? This PRD assumes one-click approval is not automated but adds G2 as defence in depth. Confirm before build.

4.4 Agent experience

  • A1. Draft appears pre-filled in the reply box, visibly marked as AI-drafted until edited.
  • A2. Discard requires a reason: wrong answer, wrong tone, missing info, wrong queue, should not be drafted.
  • A3. Agents never lose their normal tools; saved replies remain available manually.
  • A4. Customers are not told a draft was AI-assisted (agent authors the sent message). Support Ops to confirm this against Ledgerly's AI disclosure policy.

5. ML approach, data and evaluation

Data. Two years of resolved tickets (\~900k) with final queue, agent replies and saved-reply usage.

  • Label routing from the final queue, not the initial one; tickets that were reassigned are the most valuable examples.
  • Map historic replies to saved-reply IDs where possible (exact/near-match), giving a supervised "which reply fits" dataset.
  • Down-weight or exclude replies older than any policy or pricing change; Support Ops supplies the change dates. Stale answers are the main quality risk.
  • Strip PII before any use outside the production data boundary; confirm data-processing terms with Legal/DPO for customer ticket content.

Approach. Start with the simplest thing that hits the targets: an LLM classifier (or fine-tuned small model if cost/latency requires) for routing; retrieval over saved replies + resolved tickets + help centre, then LLM personalisation for drafting; a separate refund detector. Model choice is ML's call, constrained by the latency targets in §3 and D1 and per-ticket cost.

Offline evaluation (gate to pilot)

  • Held-out set: most recent 3 months, time-split (no leakage).
  • Routing: per-queue precision/recall and calibration; auto-assign coverage at the chosen threshold.
  • Drafting: saved-reply selection top-1 accuracy; 500-ticket human-graded sample scored by senior agents (correct, complete, on-policy, tone) with a ≥ 80% "send with light or no edit" bar.
  • Refund detector: recall ≥ 99% on Legal-reviewed set.
  • Factual accuracy: 0 invented account facts in the graded sample.

Online monitoring. Dashboards per queue for acceptance, edit distance, discard reasons, reassignments, FRT, CSAT and refund flags. Weekly review of a random sample of sent drafted replies by Support Ops QA.

6. Rollout, risks and open questions

Rollout

PhaseScopeExit criteria
0. Shadow (2 wks)Route + draft on all tickets, invisible to agentsOffline bars met on live traffic
1. Pilot (3 wks)Billing queue, 8 agents; routing suggest-onlyAcceptance ≥ 40%, no CSAT drop, 0 refund-guardrail misses
2. ExpandAll queues drafting; routing auto-assign above thresholdMetrics in §2 hold for 2 consecutive weeks
3. Steady stateMonthly threshold review, quarterly saved-reply refresh—

Billing goes first: highest volume, most saved-reply coverage, and it exercises the refund guardrail early under close watch.

Risks

  • Rubber-stamping. Agents approve drafts without reading them. Mitigate: placeholders block send, refund gate, QA sampling, track time-on-draft before send.
  • Stale policy in drafts. Mitigate: saved replies as primary source, policy-change exclusion in training data, owner for saved-reply freshness.
  • Misrouting Account access. Mitigate: recall-weighted threshold (R3), Triage fallback.
  • Agent trust. Early bad drafts kill adoption. Mitigate: shadow phase, drafts only above a quality-confidence score; otherwise show the suggested saved reply only.

Open questions

  1. Legal: does agent one-click approval count as "automated sending" (§4.3)?
  2. Support Ops: who can authorise refunds, and at what amounts?
  3. Eng: which helpdesk platform APIs support pre-filled internal drafts and blocking send?
  4. DPO: approval to use historic ticket content with a model provider.
  5. Support Ops: AI-assistance disclosure to customers (A4).

Owners: Eng lead (integration, guardrail enforcement, UI), ML lead (models, eval, monitoring), Support Ops lead (saved replies, policy, pilot, QA).

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisSubstantial rework

    Source: VolumesAverage first response is 7 hours against a target of under 2.

    What we’d changePresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: RiskLegal requires no automated sending of refund commitments.

    What we’d changeState the refund rule once and consistently: generated drafts never contain refund commitments, and any commitment an agent adds goes through the separate confirmation step.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changePresent it as an assumption to check with Support Ops. The brief doesn't rank the cost of misrouting by queue, and the routing design leans on it.

Claims the judge couldn’t find in the brief

  • Most of that delay is not writing time. It is triage and lookup
  • Tickets wait unassigned or in the wrong queue, then an agent spends minutes finding the right macro and filling in account details
  • Two years of resolved tickets (~900k) with final queue, agent replies and saved-reply usage are available
  • Account access is the costliest miss (locked-out customers)
  • Billing queue has the most saved-reply coverage

Grades and run details

Decision model 75 · LLM judge 8 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly48%
  • passAddresses the actual decision73%
  • passRespects explicit constraints34%
  • passIdentifies material uncertainty85%
  • failAvoids unsupported claims73%
  • passProduces the required deliverable65%
  • passSpecifies AI failure behaviour97%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs73%
  • partialProposes tests that could fail68%
Artefacts
Run
Run
#1
Time to output
2.0 min
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It likely exceeds the 1,000–1,500 word aim and its refund guardrail can still surface a draft containing refund language after stripping, so the prohibition is not cleanly enforced.

GPT-6.1 Sol · API

It respects the named audience, length, human approval requirement, and legal prohibition on automated refund commitments.

Opus 5.5 · Claude

Respects all constraints: addresses the intended readers, within word count, mandates agent approval, and enforces no automated refund commitments via guardrails.

Avoids unsupported claimsWrongRightWrong
Gemini 3.8 Flash · API

It presents several unestablished current-state assumptions as facts, such as existing channels, plan metadata, PII fields, CSAT data, and infrastructure.

GPT-6.1 Sol · API

Normative requirements and design rules are framed as proposals, not as established facts from the supplied evidence.

Opus 5.5 · Claude

Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a detailed PRD, but it is too long for the requested length and includes invented operational details that would need verification before use.

GPT-6.1 Sol · API

It is a complete PRD for engineering, ML, and support-operations leads that could start work from.

Opus 5.5 · Claude

Provides a complete PRD for the required audience, within the 1,000–1,500 word range, that they could act on.

Success settles trade-offsWrongWrongRight
Gemini 3.8 Flash · API

It sets metric targets but does not define a time frame for achieving FRT or an explicit rule for which goal wins when metrics conflict.

GPT-6.1 Sol · API

It does not give an explicit trade-off rule such as accepting lower coverage to preserve precision above a stated level.

Opus 5.5 · Claude

Sets success metrics with targets (FRT <2h, routing accuracy ≥92%) and explicit trade-off rules like prioritizing account access recall at the expense of other queues.

Proposes tests that could failWrongWrongRight
Gemini 3.8 Flash · API

Some gates have thresholds, but several lack clear measurement windows or explicit actions for failure, such as the GA FRT target and overall launch decision.

GPT-6.1 Sol · API

Several acceptance tests and kill criteria lack numeric thresholds, measurement windows, or the action each outcome triggers.

Opus 5.5 · Claude

Each pilot phase has numeric exit criteria (e.g., acceptance ≥40%, 0 refund misses), measurement windows (3 weeks), and triggers that stop rollout if not met.

All got wrong 1

Uses the supplied evidence correctlyWrongWrongWrong
Gemini 3.8 Flash · API

The PRD invents several current-state facts, including 7.1-hour FRT, manual triage details, channels, customer metadata, PII/CSAT/re-route data, and existing systems.

GPT-6.1 Sol · API

It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.

Opus 5.5 · Claude

The PRD states that the delay cause is triage and lookup and that tickets wait unassigned, and assumes saved-reply usage is in the historical data, none of which is evidence from the brief.

All got right 4

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits to building AssistRoute as a human-approved drafting and routing feature and frames the decision for engineering, ML, and support-operations leads.

GPT-6.1 Sol · API

It commits to an MVP with a clear approval flow, release gates, and conditions that would change rollout.

Opus 5.5 · Claude

Commits to building the feature with phased rollout and exit criteria that would halt expansion if not met, framed for engineering/ML/support-ops leads.

Identifies material uncertaintyRightRightRight
Gemini 3.8 Flash · API

It identifies material unknowns such as routing accuracy, draft utilization, latency, model drift, and agent over-reliance, and proposes shadow mode, canary, and gates to resolve them.

GPT-6.1 Sol · API

It names missing channels/languages, unmeasured thresholds, sample-size limitations, and how evaluation and pilot results would resolve them.

Opus 5.5 · Claude

Identifies open questions about legal definition, refund authorizations, platform APIs, data use, and AI disclosure, with owners and action to resolve.

Specifies AI failure behaviourRightRightRight
Gemini 3.8 Flash · API

It specifies low-confidence routing to General with needs_triage, low-confidence drafting suppression, and refund-violation handling.

GPT-6.1 Sol · API

It specifies low-confidence routing to manual triage, abstention or clarification when evidence is missing, stale-draft disabling, and outage fallback.

Opus 5.5 · Claude

Defines low-confidence routing to a Triage view, placeholders blocking send, draft suppression for sensitive cases, and discard reasons.

Defines evals and a launch barRightRightRight
Gemini 3.8 Flash · API

It defines offline validation, shadow mode, canary gates, routing accuracy, utilization, latency, and zero refund breach thresholds.

GPT-6.1 Sol · API

It defines an adjudicated test set, quality metrics, thresholds, and launch gates.

Opus 5.5 · Claude

Specifies offline evaluation with human-graded sample and 80% send-with-light-edits bar, routing precision/recall thresholds, and a refund-detector recall ≥99%.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD