Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given3 items: Scenario, Volumes, Risk
ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer doesThe answer key the graders mark against

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score at 40)
  • Allows automated refund commitments
Case

Staff-level · v1.11 · synthetic · AI product, support

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 10 checks passed

A PM would need to remove or verify invented current-state facts and tighten the refund guardrail and length before using it. The reasoning is otherwise strong but not safe to ship as-is.

Output

Product Requirements Document (PRD)

Feature: AssistRoute (Automated Queue Routing & First-Response Drafting) Target Delivery: Q3 Status: Ready for Engineering Review Owners: Engineering Lead, ML Lead, Support Operations Lead

---

1. Problem Statement & Business Opportunity

Ledgerly processes 9,000 support tickets per week with a team of 42 agents across four queues: Billing (38%), Technical, Account Access, and General.

  • Current FRT: 7.1 hours average first response time.
  • Target FRT: < 2.0 hours.
  • Manual Bottlenecks: Every incoming ticket requires manual reading, classification, and initial triage. Furthermore, approximately 60% of all initial inquiries map cleanly to one of ~40 canonical saved replies (macros), yet agents spend minutes manually searching, adapting, and populating them.

Solution Overview

AssistRoute is a two-part machine learning pipeline that acts on incoming tickets: 1. Predictive Routing: Automatically assigns incoming tickets to one of the four queues. 2. First-Response Drafting: Generates an editable, context-aware draft response pre-populated in the agent’s console before ticket open.

Core Invariant: Human-in-the-Loop (HITL). No draft is ever sent directly to a customer. Support agents retain 100% send authority.

---

2. Key Objectives & Metrics

Metric CategoryTargetMeasurement Method
:---:---:---
First Response Time (FRT)$\le 2.0\text{ hours}$ (across all queues)Timestamp delta: `ticket.createdat` to `firstagentmessage.sentat`.
Routing Accuracy$\ge 93\%$ overall accuracyEvaluated against tickets reassigned to a different queue within 24h.
Billing Routing Accuracy$\ge 95\%$ precision/recallSpecific tracking on Billing due to volume (38%).
Draft Utilization Rate$\ge 65\%$ of ticketsProportion of first responses where the agent accepts the draft (as-is or edited).
Draft Edit Distance$\le 30\%$ Levenshtein edit distanceMeasures generation quality on accepted drafts.
Zero Refund Violation0 instancesHard constraint: Zero drafts promising or confirming refunds reach customers without human authorization.

---

3. Scope & Non-Goals

In Scope

  • Asynchronous classification and routing pipeline for all newly created tickets via email and web form.
  • Generation of a single personalized initial draft based on historical resolutions and the 40 standard macros.
  • Guardrail pipeline strictly forbidding autonomous refund/credit commitments.
  • Feedback loop telemetry (tracking agent accepts, edits, discards, and re-routes).

Out of Scope (Phase 1)

  • Autonomous sending (auto-resolution without human click).
  • Chat/live messaging support (limited strictly to asynchronous ticketing channels).
  • Processing multi-turn responses (AssistRoute generates first responses only).
  • Multi-language support (English only).

---

4. System Architecture & Workflow

``` [Customer Submits Ticket] │ ▼ [Event Ingestion: Webhook] ──▶ [PII Masking & Sanitization] │ ┌───────────────────────┴───────────────────────┐ ▼ ▼ [Routing Classifier] [Draft Generation Pipeline] │ │ Confidence $\ge$ Threshold? │ ├── Yes ──▶ Assign Queue 1. Macro Match / Retrieval └── No ──▶ Assign "General" (Flagged) 2. Prompt Compilation (LLM) │ 3. Deterministic Policy Guardrails │ │ └───────────────────────┬───────────────────────┘ ▼ [Agent Console: Pre-populated View] ├── Action A: Accept & Send ├── Action B: Edit & Send ├── Action C: Discard Draft └── Action D: Re-assign Queue ```

End-to-End Latency SLA

From `ticket.created` webhook ingestion to draft persistence in the ticketing database: $\le 5.0$ seconds (P95).

---

5. Functional Requirements

5.1 Routing Engine (ML Service)

  • FR-1.1: The model must classify each ticket into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • FR-1.2: Confidence Scoring:
  • If `confidence >= 0.85`: Automatically set the ticket's `queue_id` attribute.
  • If `confidence < 0.85`: Set `queueid` to `General` and append the tag `needstriage`.
  • FR-1.3: The Routing Engine must evaluate and route the ticket before generating the draft, as draft prompts require queue-specific system contexts.

5.2 Retrieval & Generation Pipeline

  • FR-2.1 (Context Hydration): The generation service must pull:
  • The ticket subject and body.
  • The authenticated user's metadata: Plan tier (`Solo`, `Growth`, `Enterprise`), account age, and active add-on modules.
  • The closest semantic match from Ledgerly's 40 standard macros.
  • FR-2.2 (Draft Generation): The system must generate a friendly, concise first response that adheres to Ledgerly brand voice, incorporates the customer's name, and directly addresses the primary issue using the retrieved macro logic.
  • FR-2.3 (Confidence Gating): If the model's semantic similarity score against approved knowledge bases or historical solutions falls below `0.70`, no draft shall be rendered. The UI must show: "Draft unavailable: Low context confidence."

5.3 Legal & Compliance Guardrail (Refund Protection)

* FR-3.1 (Deterministic Regex & Semantic Filter): Every draft must pass through a two-stage financial commitment filter prior to saving: 1. Lexical Check: Negative keyword blocking: `refund`, `credit`, `reimburse`, `waive fee`, `reverse charge`, `money back`, `comp`. 2. Semantic Entailment Check: Verification that the text does not commit Ledgerly to monetary compensation or liability. * FR-3.2 (Violation Action): If a draft contains a prohibited financial commitment: * The drafting engine must automatically strip the phrase and substitute safe boilerplate: "I have escalated your billing inquiry to our finance specialists for immediate review." * The draft must display a visible amber badge in the Agent UI: "Financial Review Required — No Commitments Permitted."

5.4 Support Agent Console UX

  • FR-4.1 (Draft Presentation): The draft must render inside the native response editor as an uncommitted state (distinct background styling or watermarked).
  • FR-4.2 (Interaction Actions):
  • Tab / Single-Click Acceptance: Agent hits `Tab` or clicks `Use Draft` to copy the generated text into the active editor.
  • Direct Editing: Editing the draft directly must be seamless without extra confirmation modals.
  • Clear / Reject: Single-click action to clear the generated text and open a blank response field.
  • FR-4.3 (Metadata Panel): A sidebar panel must show:
  • Queue classification and confidence score.
  • Which of the 40 macros informed the draft.
  • A one-click queue reassignment dropdown.

``` +-------------------------------------------------------------------------+ | Ticket #10492: "Charged twice for Payroll add-on" | | Sender: Sarah Jenkins (Growth Tier) | Queue: [Billing (96%)] [Re-route] | +-------------------------------------------------------------------------+ | [ AI Drafted Response ] [Macro Source: Billing_Dup]| | | | Hi Sarah, | | | | Thanks for reaching out. I see that your account was debited twice | | for the Payroll add-on on October 1st. | | | | I have escalated this to our payment operations team to review the | | duplicate transaction details. We will update you here as soon as this | | review is complete. | | | | [ Accept Draft (Tab) ] [ Edit Draft ] [ Reject ] | +-------------------------------------------------------------------------+ | [ Agent Response Box ] | +-------------------------------------------------------------------------+ ```

---

6. ML & Data Engineering Specifications

6.1 Data Preparation (2-Year Historical Corpus)

1. Sanitization: Strip all historical PII (tax identifiers, SSNs, credit card numbers, passwords) using Microsoft Presidio or an equivalent NER pipeline before training/indexing. 2. Filtering: * Drop all historical tickets that required more than 4 re-routes (noisy labels). * Filter out tickets closed with negative customer satisfaction (CSAT $\le 2$). * Exclude responses superseded by outdated accounting rules or old pricing tiers (Ops team to define date cutoffs). 3. Macro Ground-Truth Alignment: Map the 40 canonical macros against historical agent responses to serve as gold-standard reference pairs.

6.2 Model Specifications

  • Routing Classifier:
  • Architecture: Fine-tuned lightweight encoder (e.g., `modern-bert-base` or `RoBERTa-base`) or an optimized classification endpoint.
  • Input: Ticket Subject + Body.
  • Output: Softmax distribution over `[Billing, Technical, Account Access, General]`.
  • Latency Target: $< 200\text{ ms}$.
  • Generative Drafting Model:
  • Architecture: Hosted LLM (e.g., Claude 3.5 Sonnet or GPT-4o-mini) via secure enterprise API with zero-data-retention agreements.
  • Prompt Design: System prompt containing Ledgerly tone guidelines, user context JSON, the selected macro instructions, and strict instructions forbidding financial promises.
  • Temperature: `0.1` (low variability, high determinism).

6.3 Telemetry & Event Logging

The frontend and backend must emit the following events to the analytics lakehouse: * `ticketrouted`: `{ ticketid, predictedqueue, confidence, autoassigned: bool }` * `ticketrerouted`: `{ ticketid, oldqueue, newqueue, agentid }` * `draftgenerated`: `{ ticketid, macroid, promptversion, modelid, generationtimems }` * `draftactioned`: * `action`: `ACCEPTEDASIS` | `EDITED` | `REJECTED` * `originaldraft`: string * `finalsenttext`: string * `levenshteindistance`: int * `agentid`: string

---

7. Safety, Compliance, and Security

  1. Refund Protection Hard Filter: As specified in Section 5.3, any draft containing automated refund promises is blocked from the UI by an air-gapped deterministic validation check that runs after LLM generation and cannot be bypassed by prompt engineering.
  2. Zero Training on Customer Data: LLM vendors must have explicit Data Processing Agreements (DPAs) confirming customer ticket content is not used to train foundation models.
  3. Data Storage & SOC 2: Masked inputs and outputs stored for model evaluation must adhere to standard data retention policies (auto-purged after 90 days).

---

8. Rollout Plan & Milestones

``` Phase 0: Offline Validation & Guardrail Auditing (Weeks 1-3) ├── Train routing model on historical data; evaluate against held-out test set. ├── Run 10,000 synthetic adversarial prompts targeting the refund guardrail. └── Success Gate: Routing accuracy > 93%, 0 adversarial refund breaches.

Phase 1: Shadow Mode (Weeks 4-5) ├── Pipeline processes live tickets; outputs are logged to Databricks/Snowflake. ├── No UI visible to agents. └── Evaluate: Model inference latency, real-time routing accuracy vs. manual triage.

Phase 2: Single-Queue Canary Pilot (Weeks 6-7) ├── Deploy exclusively to the Billing Queue (10 designated agents). ├── Daily sync with Support Operations leads to review draft quality. └── Success Gate: Draft utilization > 50%, no severe agent blockers.

Phase 3: General Availability (GA) & Iteration (Weeks 8+) ├── Enable routing and drafting across all 42 agents and 4 queues. ├── Implement weekly retraining cadence for classifier based on override logs. └── Track progress toward < 2.0 hour FRT target. ```

---

9. Risks & Mitigations

RiskImpactLikelihoodMitigation Strategy
:---:---:---:---
Agent over-reliance (rubber-stamping bad drafts)HighMediumRandom 2% QA audits on accepted drafts by Support Operations leads. Prominent UI reminders that agents are solely responsible for sent text.
Hallucinated accounting/tax guidanceCriticalLowRestrict generation strictly to standard macros and product knowledge base; temperature set to 0.1; fallback to "Draft unavailable" when query context is ambiguous.
Model drift due to product updatesMediumHighWeekly regression testing on macro sets. Support Ops manages a macro repository that directly updates retrieval indexes without code deploys.
Routing loop (agent re-routes ticket repeatedly)LowLowHard ceiling: If a ticket is manually re-routed twice, AssistRoute disables auto-routing for that ticket and locks it to manual handling.

Claims the judge couldn’t find in the brief

  • Current average first response time is 7.1 hours.
  • Every incoming ticket requires manual reading, classification, and initial triage.
  • Agents spend minutes manually searching, adapting, and populating saved replies.
  • New tickets arrive via email and web form.
  • Authenticated user metadata includes plan tier (Solo, Growth, Enterprise), account age, and active add-on modules.
  • Historical tickets contain PII such as tax identifiers, SSNs, credit card numbers, and passwords.
  • Historical tickets include CSAT scores and re-route counts.
  • The company uses Databricks/Snowflake and has SOC 2 controls.

Grades and run details

Decision model 65 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly59%
  • passAddresses the actual decision87%
  • partialRespects explicit constraints26%
  • partialIdentifies material uncertainty34%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable55%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar92%
  • partialSuccess settles trade-offs75%
  • partialProposes tests that could fail69%
Run
Run
#1
API response time
33 s
Submitted
1 Oct 2026

Needs more than a quick edit4 of 10 checks passed

Contains an unsupported current routing accuracy figure, missing evaluation plan/launch bar, and no trade‑off rules; a PM would need to re‑check evidence and add substantial missing sections.

Output

Product Requirements Document (PRD): AI-Powered Ticket Triage and Response Draft Generation ("AutoDraft & Route")

---

1. Document Overview

MetadataDetails
:---:---
Project NameAutoDraft & Route (Ledgerly Support Intelligence)
Target AudienceEngineering, Machine Learning (ML), and Support Operations (Ops) Leads
Document StatusDraft / Ready for Engineering Review
Target LaunchQ1 2027
GoalReduce Average First Response Time (FRT) from 7 hours to <2 hours while maintaining human-in-the-loop safety and strict legal compliance on financial commitments.

---

2. Background & Problem Statement

2.1 Current State & Metrics

Ledgerly is a premier accounting software platform designed specifically for small businesses. As our customer base scales, our support infrastructure is facing severe volume pressures: * Support Team Size: 42 agents distributed across four specialized queues: Billing, Technical, Account Access, and General. * Ticket Volume: ~9,000 incoming tickets per week, with 38% (~3,420 tickets/week) falling into the Billing queue. * First Response Time (FRT): Current average is 7 hours, severely lagging behind our target of under 2 hours. * Repetitive Workflows: Approximately 60% of all incoming tickets are resolved using one of roughly 40 standard saved replies (e.g., password resets, invoice retrieval, subscription tier changes).

2.2 The Opportunity

We possess two years of historical resolved tickets, including full conversation histories and agent replies. This corpus serves as a robust training and evaluation dataset for an ML-driven system that can: 1. Accurately classify and route incoming tickets to the correct queue. 2. Automatically draft context-aware first responses using our proven saved replies and historical resolutions.

2.3 Core Constraint (Legal Compliance)

  • Zero Automated Sending: Legal has mandated a strict human-in-the-loop requirement. The AI system must never autonomously send messages to customers.
  • Financial Safety: The model is strictly prohibited from making automated commitments regarding refunds, credits, or billing adjustments without explicit human review and approval.

---

3. Goals and Success Metrics (KPIs)

MetricCurrent BaselineTarget (3 Months Post-Launch)
:---:---:---
Average First Response Time (FRT)7 hours< 2 hours
Routing Accuracy~75% (Manual/Rule-based)1> 95% across all four queues
Draft Acceptance RateN/A> 70% (used as-is or with minor edits)
Agent Handle Time (AHT)Baseline-25% reduction in time spent crafting initial replies
Compliance / Unauthorized Sends00 violations (Hard safety guardrail)

---

4. User Personas & Workflows

4.1 Support Agent

  • Role: Resolves customer queries, reviews AI drafts, edits when necessary, approves responses, and handles complex escalations.
  • Workflow: Opens the ticketing dashboard, sees incoming tickets pre-sorted into their designated queue with an AI-generated draft response attached. The agent reviews the draft, clicks "Approve & Send" or modifies the text, ensuring zero financial compliance risk.

4.2 Support Operations Lead

  • Role: Manages queue health, tracks SLAs, audits agent performance, and maintains the library of saved replies and routing rules.
  • Workflow: Monitors routing precision dashboards, updates macro/saved reply templates, and analyzes draft acceptance/rejection feedback loops to refine prompts.

4.3 Engineering & ML Leads

  • Role: Build, deploy, scale, and monitor the inference pipeline, classification models, LLM orchestration, and UI integration.

---

5. Functional Requirements

5.1 Automated Ticket Routing (Triage Engine)

  • Queue Classification: Upon ticket ingestion, the ML classification service must categorize incoming text into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • Confidence Thresholds:
  • Confidence $\ge$ 85%: Automatically assign the ticket to the predicted queue.
  • Confidence < 85%: Route the ticket to a "Triage Review" sub-queue or default General queue with a visual tag flagging low classification confidence.
  • Metadata Tagging: Tag tickets with sub-intent labels (e.g., `Billing: Refund Request`, `Technical: Bank Sync Failure`) to assist the response generation engine.

5.2 Response Draft Generation

  • Context Window & Retrieval: The generation engine must leverage the incoming user query, customer account metadata (plan tier, active status), and relevant historical examples from the 2-year resolved ticket dataset (via RAG / vector search).
  • Saved Reply Integration: The model must prioritize mapping inquiries to Ledgerly's 40 core saved replies where applicable, maintaining brand voice, accuracy, and tone consistency.
  • Draft Status: Every generated response must be marked with a distinct internal status: `Draft - Pending Agent Review`. It must remain locked from customer view until explicit human sign-off.

5.3 Agent Workspace UI/UX Integration

  • Side-by-Side Review: The agent UI must display the incoming customer message alongside the AI-generated draft in a clear, editable text box.
  • Action Controls:
  • [Approve & Send]: Immediately transmits the draft to the customer and logs the interaction.
  • [Edit & Send]: Allows agents to modify text inline before sending. Edits must be logged for ML fine-tuning feedback loops.
  • [Discard & Rewrite]: Clears the draft if the AI misunderstood the prompt.
  • Visual Safety Badges: Clear UI warnings must appear on any ticket flagged as containing billing or refund keywords, reminding agents of compliance protocols.

---

6. Non-Functional & Safety Requirements

6.1 Legal & Financial Guardrails (Critical)

  • No Auto-Dispatch: Zero API endpoints or automation rules are permitted to transition a draft state to `Sent` without a cryptographic token or database action originating from an authenticated human agent session.
  • Refund Keyword Interceptor: A deterministic rule-based filter must scan all generated drafts for high-risk financial terminology (`refund`, `credit`, `waiver`, `chargeback`). If detected, the UI must highlight the section in amber and require a secondary manual checkbox confirmation from the agent before sending.

6.2 Performance & Latency

  • Inference Latency: End-to-end classification and draft generation must complete within < 3.0 seconds of ticket ingestion via webhook, ensuring drafts are ready before an agent opens the ticket.
  • Availability: 99.9% uptime SLA for the ML inference and routing service.

6.3 Data Privacy & Security

  • PII Redaction: Customer tax IDs, banking credentials, and sensitive financial figures must be handled in compliance with GDPR, CCPA, and Ledgerly security policies. Training data pipelines must scrub extraneous PII prior to embedding generation.

---

7. Data, ML, & Training Strategy

7.1 Dataset Utilization

  • Corpus: 2 years of historical resolved tickets (approx. 900,000+ interaction pairs across 42 agents).
  • Data Splitting: 80% training/fine-tuning, 10% validation, 10% held-out test set reflecting recent seasonal billing spikes.

7.2 Model Architecture & Approach

  • Classification: Fine-tuned encoder model (e.g., BERT/DeBERTa variant) for multi-class queue routing and intent tagging.
  • Generation: Instruction-tuned LLM accessed via secure API or private VPC deployment, augmented with a vector database (RAG) containing the 40 approved saved replies and top-performing historical resolutions.
  • Feedback Loop: Log all agent edits (diffs) and rejections to construct a weekly RLHF (Reinforcement Learning from Human Feedback) or supervised fine-tuning dataset to reduce hallucination rates.

---

8. High-Level Architecture & Data Flow

``` [Customer Email / Web Form] │ ▼ [Ingestion Service / Webhook] │ ├──► [ML Classification Engine] ──► Assigns Queue & Sub-intent │ └──► [RAG Vector DB (Saved Replies & History)] │ ▼ [LLM Draft Generation Service] ──► Generates Initial Response Text │ ▼ (Strictly Locked: Status = Pending Review) [Support Ticketing Database] │ ▼ [Agent UI Dashboard] ◄─── Human Agent Reviews, Edits, & Approves │ ▼ (Explicit Human Action) [Customer Dispatch API] ```

---

9. Phased Rollout Plan

  • Phase 1: Shadow Mode (Weeks 1–4)
  • Deploy classification and draft generation in the background.
  • Log routing accuracy and draft quality without surfacing drafts to agents. Establish baseline ML accuracy metrics.
  • Phase 2: Internal Pilot / Beta (Weeks 5–8)
  • Roll out the UI feature to a pilot group of 5 senior Billing agents.
  • Measure acceptance rates, time saved, and friction points. Refine prompt engineering and guardrails.
  • Phase 3: General Availability Across All Queues (Weeks 9–12)2
  • Roll out to all 42 agents across Billing, Technical, Account Access, and General queues.
  • Establish continuous monitoring dashboards for Support Ops and Engineering leads.

---

10. Open Questions & Risks

  1. Edge Cases in Billing: How should the model handle complex multi-invoice dispute threads where historical context spans multiple months? (Mitigation: Surface the last 3 ticket summaries alongside the draft).
  2. Agent Adoption: Will agents trust the AI drafts, or will they rewrite them from scratch? (Mitigation: Emphasize time savings during training and incorporate agent feedback buttons).

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimStart again

    Source: Volumes9,000 tickets a week; 38% are billing.

    What we’d changeRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.

  2. 2Test or gate too weakTighten the testSubstantial rework

    What we’d changeAdvance each phase on quality gates (routing precision, draft acceptance, zero refund-commitment misses) rather than the calendar, with a human grading rubric and a rollback trigger.

Claims the judge couldn’t find in the brief

  • Current routing accuracy is ~75% (Manual/Rule-based).

Grades and run details

Decision model 55 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly63%
  • passAddresses the actual decision47%
  • partialRespects explicit constraints41%
  • partialIdentifies material uncertainty90%
  • partialAvoids unsupported claims48%
  • partialProduces the required deliverable43%
  • passSpecifies AI failure behaviour26%
  • partialDefines evals and a launch bar84%
  • partialSuccess settles trade-offs82%
  • partialProposes tests that could fail74%
Artefacts
Run
Run
#1
Time to output
14 s
Submitted
24 Sept 2026

Usable with a quick edit10 of 10 checks passed

The PRD is thorough, actionable, and requires at most minor wording edits before engineering, ML, and support-ops leads can start work.

Output

Ledgerly: AI first-response drafts and ticket routing

Status: Proposed MVP requirements. Numerical gates below are proposed targets, not observed performance. Product owns scope; Engineering owns workflow integrity; ML owns model quality; Support Operations owns queue policy, content and review quality.

1. Problem and outcome

Ledgerly’s 42 support agents handle 9,000 tickets weekly across Billing, Technical, Account access and General. Billing represents 38%, approximately 3,420 tickets weekly. Average first response is seven hours against a target below two. Approximately 60% of tickets use one of 40 saved replies.

The feature should reduce triage and composition time without shifting effort into correcting poor drafts or recovering misrouted tickets. Every outgoing draft requires human approval. Legal prohibits automated sending of refund commitments.

Faster drafting alone may not close the five-hour gap. Before piloting, Support Operations must measure arrival-to-assignment, assignment-to-review and review-to-send delays by queue and staffed hours, then identify any coverage or staffing changes needed.

2. Scope and operating boundary

MVP processes newly created English-language, text-based tickets in the existing support workspace. It assigns a queue and prepares a first-response draft. Threads, ticket metadata and approved support content are inputs. Attachments are not interpreted; attachment-dependent or unsupported-language tickets remain available for manual handling.

Excluded: follow-up generation, autonomous sending, refunds, account changes, payment actions and staffing optimisation. Existing spam, security and priority rules run first and cannot be overridden by the model.2

Routing and drafting operate independently. A drafting failure must not prevent routing or agent access. An uncertain route must not prevent a useful draft.

3. Agent workflow

  1. Ticket creation starts asynchronous processing. The ticket is immediately visible and its response clock starts at original receipt.
  2. The system records a route, reason and confidence tier, then prepares a draft where supported.
  3. The agent sees the original message, assigned queue, editable draft, source references and warnings. Sources and warnings are internal only.
  4. The agent can change queue, edit, regenerate, discard or write manually. Regeneration never overwrites unsaved edits without confirmation.
  5. Selecting Approve and send authorises the exact visible text. There is no bulk approval.

Draft states are pending, ready, unavailable, stale and sent. Queue changes preserve receipt time and draft history. New customer messages or changes to relevant ticket context mark drafts stale and require renewed review. Show actionable failure messages and keep the manual composer available.

4. Routing requirements

Support Operations owns this initial taxonomy:

QueuePrimary issue
BillingCharges, subscriptions, invoices, cancellations and refund requests
TechnicalErrors, integrations, imports and malfunctioning features
Account accessLogin, authentication, permissions and suspected account takeover
GeneralProduct guidance and genuinely uncategorised enquiries

For multiple intents, suspected account compromise takes precedence, followed by access-blocking issues; otherwise route by the customer’s main requested resolution. Record secondary intents for the receiving agent.

Automatically assign only when a queue-specific threshold meets the evaluation gate below. Confidence must be calibrated against labelled examples, not taken from a model’s self-reported certainty.

Below threshold, assign to General with a distinct Needs triage status and show the leading suggestions. This is a fallback assignment, not a successful classification. Support Operations assigns a named triage owner each shift and reviews these tickets at least every 30 minutes during staffed hours. Existing out-of-hours escalation remains in force.

Agents can override any route and optionally record a reason. Never automatically reroute after an agent takes ownership. Log overrides for review, not immediate retraining.

5. Draft content and refund controls

Start with retrieval from the approximately 40 saved replies and current, Support Operations-approved help and policy content. Adapt an applicable reply before attempting a novel answer. Each source has an owner, version and review date; withdrawn content becomes unavailable immediately.

Drafts must answer the stated question, request necessary missing details and cite supporting sources internally. They must not invent account facts, troubleshooting outcomes, eligibility, amounts or dates. Account-specific claims require authorised, current account context. Without it, draft a clarification or indicate that an agent must investigate.

For refund requests, MVP drafts may acknowledge the request and explain approved review steps, but must not promise eligibility, payment amounts or payment dates. Agents may manually add commitments only under Ledgerly’s existing refund authority policy. Flag refund-related tickets and require an explicit acknowledgement when the final text contains a detected commitment. This check assists reviewers; it is not the legal enforcement boundary.

The enforcement boundary is the sending service: generation workers have no send credentials. Every AI-assisted send requires an authenticated agent approval bound to ticket ID, recipient and exact message version. Any subsequent edit invalidates approval. Retries cannot send duplicate messages. Existing automation must not consume AI drafts as sendable replies. Therefore, missed refund detection cannot trigger autonomous sending.

6. Data and ML approach

Use the two-year resolved-ticket archive to learn routing patterns and evaluate drafting. Historical agent replies are examples, not authoritative policy. Resolution does not establish correctness or refund permission.

Support Operations must relabel a representative sample against today’s taxonomy, with two reviewers and adjudication for disagreements. Redact credentials, payment details and unnecessary personal information. Preserve tenant boundaries and restrict access to approved training personnel and services.

Split chronologically into training, validation and a locked recent test set. Keep complete conversations and duplicate ticket clusters in one split. At inference and evaluation, expose only information available before the first response; later replies and final queue labels must not leak into inputs.

Begin with a classifier plus retrieval-grounded generation. Fine-tuning is optional and requires measurable improvement over this baseline. No automatic learning from approvals. Provider data use and retention must be approved before production data leaves Ledgerly.

Treat customer text and retrieved content as untrusted data: embedded instructions cannot change policies, access other tenants or invoke privileged actions.

7. Evaluation and launch gates

Build a locked test set of at least 1,000 tickets, stratified across queues, with at least 150 per queue. Include saved-reply matches, ambiguous intents, refunds, outdated policies, missing context, prompt injection and sensitive-data cases. Report volume-weighted results and per-queue results separately.

Required gates:

  • Routing: at least 95% precision among automatically assigned tickets in each queue. Report confidence intervals, recall, coverage and the confusion matrix. Target at least 60% overall automatic-routing coverage; do not lower precision to achieve coverage.
  • Drafts: at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.1 Score factual accuracy, policy compliance, completeness and tone separately.
  • Critical errors: zero observed cross-tenant disclosures, unsupported refund commitments in generated drafts or account-security instructions that bypass verification. Any occurrence blocks release pending correction and regression testing.
  • Workflow: all permission, stale-approval, recipient-change, duplicate-event and retry tests pass. Attempts to send without valid approval must fail.

Use blinded human review and adjudicate disagreements. Automated grading can assist sampling but cannot determine safety gates alone. Zero observed errors is a test result, not proof of zero production risk.

8. Pilot, measurement and operations

Run one week in shadow mode across all queues, followed by a two-week pilot with agents from every queue. Randomise eligible tickets within queue and shift between assisted and existing workflows; account for shared-agent effects when interpreting results.

Primary pilot gate: at least 20% lower mean receipt-to-first-human-response time versus control. Track progress towards the under-two-hour target across all incoming tickets, including unsupported and fallback cases. Also report median, p90, percentage answered within two hours, review time, queue transfers, correction severity, backlog age and customer satisfaction. Draft creation is not a first response.

Expand only if offline gates hold, human handling time improves and transfer rates and customer satisfaction show no material deterioration. Define tolerances and adequate sample sizes before the pilot; extend measurement when inconclusive.

Engineering provides independent routing and drafting kill switches, audit logs and alerts. Target p95 draft readiness within 60 seconds; after timeout, mark unavailable and retain manual handling. Retries must not overwrite agent work.

Support Operations samples 50 assisted tickets weekly, oversampling refunds and overrides. ML monitors quality and drift by queue and content version. Any unauthorised send or cross-tenant disclosure triggers immediate suspension of the affected feature and incident review.

Before pilot launch, Engineering verifies send-path enforcement and integration contracts; ML publishes evaluation results; Support Operations approves taxonomy, sources and shift ownership; Legal confirms the refund workflow.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.

  2. 2Invented evidenceVerify or remove the claimQuick edit

    What we’d changeThe brief doesn't mention existing spam, security or priority rules, or out-of-hours escalation. Confirm they exist before the design relies on them.

Grades and run details

Decision model 90 · LLM judge 11 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly3%
  • passAddresses the actual decision82%
  • passRespects explicit constraints49%
  • passIdentifies material uncertainty83%
  • passAvoids unsupported claims75%
  • passProduces the required deliverable78%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • passSuccess settles trade-offs96%
  • partialProposes tests that could fail73%
Run
Run
#1
Time to output
1.8 min
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

The PRD invents several current-state facts, including 7.1-hour FRT, manual triage details, channels, customer metadata, PII/CSAT/re-route data, and existing systems.

Gemini 3.5 Flash-Lite · Gemini

States current routing accuracy of ~75% as fact without any support from the supplied evidence and does not label it as an assumption.

GPT-6 Astra · ChatGPT

All statements about the current situation are taken directly from the supplied context or arithmetic, with no invention.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It likely exceeds the 1,000–1,500 word aim and its refund guardrail can still surface a draft containing refund language after stripping, so the prohibition is not cleanly enforced.

Gemini 3.5 Flash-Lite · Gemini

The document remains within the 1,000–1,500‑word range, addresses the required readers, and enforces legal constraints with human‑only dispatch and a refund‑keyword interceptor.

GPT-6 Astra · ChatGPT

The no-automated-refunds requirement is enforced via draft rules, approval checks, and a send-path enforcement boundary; length and reader constraints are met.

Identifies material uncertaintyRightWrongRight
Gemini 3.8 Flash · API

It identifies material unknowns such as routing accuracy, draft utilization, latency, model drift, and agent over-reliance, and proposes shadow mode, canary, and gates to resolve them.

Gemini 3.5 Flash-Lite · Gemini

Open questions are listed but are not tied to decision‑changing thresholds or explicit resolution plans; the PRD does not state what would halt or alter the rollout.

GPT-6 Astra · ChatGPT

Unknowns like actual delay components, zero-error test limits, and pilot inconclusiveness are named, with resolution steps specified.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

It presents several unestablished current-state assumptions as facts, such as existing channels, plan metadata, PII fields, CSAT data, and infrastructure.

Gemini 3.5 Flash-Lite · Gemini

The routing accuracy baseline of ~75% is presented as a fact without being marked as an assumption, and no evidence supports it.

GPT-6 Astra · ChatGPT

Hypotheses (e.g., 'faster drafting alone may not close the gap') are clearly flagged as possibilities, not fact.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a detailed PRD, but it is too long for the requested length and includes invented operational details that would need verification before use.

Gemini 3.5 Flash-Lite · Gemini

The output is a complete PRD with functional requirements, architecture, and rollout plan that the named leads could start work from.

GPT-6 Astra · ChatGPT

A complete PRD with scope, workflow, ML approach, evals, and pilot plan is delivered for the target leads, within ~1,200 words.

Defines evals and a launch barRightWrongRight
Gemini 3.8 Flash · API

It defines offline validation, shadow mode, canary gates, routing accuracy, utilization, latency, and zero refund breach thresholds.

Gemini 3.5 Flash-Lite · Gemini

No evaluation dataset criteria, metrics (e.g., draft quality), or specific launch threshold for progressing from shadow mode to pilot or GA are defined.

GPT-6 Astra · ChatGPT

A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.

Success settles trade-offsWrongWrongRight
Gemini 3.8 Flash · API

It sets metric targets but does not define a time frame for achieving FRT or an explicit rule for which goal wins when metrics conflict.

Gemini 3.5 Flash-Lite · Gemini

Success metrics include targets and a timeframe, but no explicit trade‑off rule (e.g., coverage vs. precision) is stated.

GPT-6 Astra · ChatGPT

Success metric (≥20% reduction in response time) with a pilot time frame, and an explicit precision-over-coverage trade-off rule are given.

Proposes tests that could failWrongWrongRight
Gemini 3.8 Flash · API

Some gates have thresholds, but several lack clear measurement windows or explicit actions for failure, such as the GA FRT target and overall launch decision.

Gemini 3.5 Flash-Lite · Gemini

The phased rollout lacks numeric thresholds, measurement windows, and actions tied to results; it does not define kill criteria.

GPT-6 Astra · ChatGPT

All evaluation and pilot gates have numeric thresholds, a two-week measurement window, and defined actions (block release, suspend feature).

All got right 2

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits to building AssistRoute as a human-approved drafting and routing feature and frames the decision for engineering, ML, and support-operations leads.

Gemini 3.5 Flash-Lite · Gemini

The PRD commits to a clear design for the AI feature, framed for the engineering, ML, and ops leads.

GPT-6 Astra · ChatGPT

The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.

Specifies AI failure behaviourRightRightRight
Gemini 3.8 Flash · API

It specifies low-confidence routing to General with needs_triage, low-confidence drafting suppression, and refund-violation handling.

Gemini 3.5 Flash-Lite · Gemini

Low-confidence classification routes to a triage queue, and a deterministic refund‑keyword interceptor triggers a secondary manual checkbox.

GPT-6 Astra · ChatGPT

Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD