Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Specifies AI failure behaviour100% pass
    Fallback routing to General with Needs triage, stale draft handling, manual composer availability, and clear error messages are specified.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  2. Addresses the actual decision96% pass
    The PRD commits to an explicit MVP scope and states specific expansion conditions based on pilot results, framed for the named leads.
    GPT-6 Astra · ChatGPT · AI triage for support tickets
  3. Defines evals and a launch bar89% pass
    A 1,000-ticket stratified test set, per-queue precision/recall/coverage gates, draft sendability, and zero critical-error thresholds are defined.
    GPT-6 Astra · ChatGPT · AI triage for support tickets

Where it slips

  1. Uses the supplied evidence correctly27% pass
    It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.
    GPT-6.1 Sol · API · AI triage for support tickets
  2. Proposes tests that could fail36% pass
    Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.
    GPT-6 Luna · API · AI triage for support tickets
  3. Avoids unsupported claims39% pass
    Presents interpretations about the root cause of first-response delay and data availability (saved-reply usage, final queue labels) as fact without support.
    Opus 5.5 · Claude · AI triage for support tickets

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes each ticket to the right queue. Support agents approve every draft before it is sent. The PRD is for the engineering, ML and support-operations leads who will build and run the feature: write it so they could start work from it. Aim for 1,000–1,500 words.

What the model was given3 items: Scenario, Volumes, Risk
ScenarioLedgerly is accounting software for small businesses. The support team is 42 agents across four queues: Billing, Technical, Account access and General. Two years of resolved tickets, with the agents' replies, are available for training and evaluation.
Volumes9,000 tickets a week; 38% are billing. Average first response is 7 hours against a target of under 2. About 60% of tickets are answered with one of roughly 40 saved replies.
RiskLegal requires no automated sending of refund commitments.
What a strong answer doesThe answer key the graders mark against

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score at 40)
  • Allows automated refund commitments
Case

Staff-level · v1.11 · synthetic · AI product, support

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit7 of 10 checks passed

It is largely usable but invents current systems not in the brief, so a PM would need to verify the evidence before acting. It also lacks an explicit trade-off rule and complete test windows/actions, requiring more than a quick edit.

Output

PRD: AI-Assisted First Response and Ticket Routing

Product: Ledgerly Support Assist Audience: Engineering, ML and Support Operations leads Status: Proposed MVP; thresholds below are launch targets, not measured results Accountable owner: Support Operations lead, with Engineering and ML leads owning delivery and model quality

1. Problem and objective

Ledgerly’s 42 support agents handle 9,000 tickets per week across Billing, Technical, Account access and General. Billing represents 38% of volume—approximately 3,420 tickets weekly. Average first-response time is seven hours, against a target of under two hours. Approximately 60% of tickets receive one of roughly 40 saved replies.

The feature will classify new tickets into the appropriate queue and prepare a grounded first-response draft for an agent to review. It should reduce queue-selection work and repetitive writing without delegating customer communication or financial commitments to the model.

Every outbound response requires explicit agent approval. Legal prohibits automated sending of refund commitments. The MVP will have no automated-send path and will not generate refund promises; agents must author any commitment using the existing authorized refund process.

2. Goals and success measures

Primary goals:

  • Reduce average time to the first substantive human-approved response to under two hours during a staffed pilot.
  • Reduce median active agent time spent preparing first responses by at least 30%.
  • Route tickets accurately while avoiding hidden misroutes.
  • Preserve response correctness, account security and customer satisfaction.

Measure first-response time from ticket creation to the first substantive response sent by an agent. Automated receipts do not count. Report calendar-hour and staffed-hour results separately, plus median and p90, to prevent averages hiding long waits.

Additional proposed launch and pilot thresholds:

MeasureTarget
Queue-classification accuracy≥95% on the adjudicated holdout
Account access routing recall≥98%
High-confidence automatic-routing precision≥98%
Drafts usable with no substantive correction≥80% in blinded review
Unsupported financial or security commitmentsZero observed in launch evaluation
Draft availability latencyp95 ≤15 seconds after ingestion
Pilot qualityNo material deterioration in QA scores or customer satisfaction

Zero observed errors is a release gate, not proof of zero production risk. ML must report sample sizes and confidence intervals. Support Operations must confirm staffing coverage: drafting improvements alone cannot guarantee the response-time target.

3. Scope

MVP includes:

  • Classification into Billing, Technical, Account access or General.
  • Automatic routing only above validated, queue-specific confidence thresholds.
  • Retrieval of approved saved replies and current support documentation.
  • A first-response draft with internal source references, risk flags and suggested clarifying questions.
  • Agent controls to edit, discard, regenerate, change queue and approve/send.
  • Audit logs, monitoring and immediate disable controls.

Excluded: automatic sending, subsequent-turn assistance, ticket resolution, refunds, account changes, security verification decisions and customer-facing AI chat.

Ticket channels, languages and attachment formats are not specified. Support Operations will inventory them before implementation. MVP supports only validated text channels and languages; unsupported inputs receive normal human triage without a generated draft.

4. User workflow and functional requirements

  1. Ingest: On creation of a new eligible ticket, capture its text, subject and approved metadata. Use only authorized customer/account context available to the assigned support role.
  2. Assess: Predict a queue, calibrated confidence and risk flags such as refund request, account compromise or insufficient information.
  3. Route: Assign high-confidence tickets to the predicted queue. Send uncertain cases to the existing manual-triage destination, provisionally General. Surface uncertainty prominently; do not treat General as a confident classification.
  4. Draft: Retrieve relevant approved material and produce a concise response addressing the request. Prefer adapting an applicable saved reply over open-ended generation.
  5. Review: Show the draft, suggested queue, risk flags and source links in the agent workspace. Clearly label the text “AI draft—not sent.”
  6. Approve/send: The existing send action requires an explicit authenticated agent action. No background job, timeout or model output may trigger sending.
  7. Learn: Record queue corrections, edits, rejection reasons and quality reviews for controlled evaluation and future retraining.

Drafts must not invent transactions, troubleshooting outcomes, refund eligibility or account status. Where evidence is missing, ask for necessary information or acknowledge that an agent must investigate. Avoid requesting passwords, full payment-card details or other unnecessary sensitive data.

Refund-related drafts may acknowledge the request and explain approved next steps, but must not promise an amount, eligibility or processing date. Flag these tickets for Billing review. Agents may add commitments only after authorized verification.

A draft becomes stale if relevant ticket content or account context changes. Disable approval until it is refreshed or explicitly reviewed against the updated context. Manual queue changes take precedence over later model results.

5. Data and ML approach

Two years of resolved tickets and agent replies are available. Historical replies are examples, not authoritative policy: outdated guidance and unauthorized commitments must not be reproduced.

ML and Support Operations will:

  • Define queue labels and rules for mixed-intent tickets. Account-security concerns take precedence over routine billing questions; Support Operations must approve the complete precedence matrix.
  • Audit final queue labels and sample ambiguous cases for expert adjudication.
  • Remove duplicates, signatures and irrelevant quoted history; redact unnecessary personal and payment information.
  • Preserve ticket/thread/customer grouping across splits to reduce leakage.
  • Use a chronological training, validation and held-out test split, with the newest period reserved for testing.
  • Evaluate only information available when the ticket arrived. Later replies and resolutions may provide labels but must never become runtime inputs.

Benchmark saved-reply retrieval and a conventional classifier before introducing more complex models. Select the simplest approach meeting quality, latency and operational requirements.

Generation will use current, versioned approved content. Historical replies may support offline training subject to privacy and quality approval, but must not serve as an unrestricted runtime knowledge source. Conflicting or missing sources trigger a clarification or human-investigation draft rather than a guessed answer.

Customer text and retrieved content are untrusted inputs. Instructions embedded in tickets must not change system rules, authorize actions or bypass review.

6. Engineering design and controls

Implement an asynchronous pipeline behind feature flags:

Ticket event → eligibility check → classifier → routing decision → retrieval → generation → policy validation → draft storage → agent UI.

Persist ticket ID/version, model and prompt versions, retrieved document versions, queue scores, draft state, risk flags and agent actions. Use idempotency keys to prevent duplicate processing and concurrency controls to avoid overwriting agent work.

The AI service must have no credential or permission to send messages, issue refunds or modify accounts. Routing permissions must be limited to approved queues. Outbound messages remain controlled by the existing authenticated support application.

Post-generation checks will block prohibited promises and unsupported sensitive claims. Failed checks suppress the draft and display a reason; they must not merely append a disclaimer.

Apply role-based access, encryption and existing retention rules. Any external model provider requires Security and Legal approval, including contractual restrictions on retention and training use. Log identifiers and operational metadata where possible, not unrestricted ticket bodies.

On timeout, provider outage or validation failure, preserve normal ticket handling. Keep the ticket visible, apply manual triage where necessary and show “Draft unavailable.” Never delay ticket intake while waiting for AI.

7. Evaluation and release gates

Build an adjudicated test set covering all queues, common saved-reply cases, mixed intent, sparse descriptions, refund disputes, account compromise, outdated-policy examples and prompt-injection attempts. Report overall and per-queue performance; oversampled risk cases must also be reported separately from production-weighted results.

Two support reviewers will score drafts for correctness, relevance, completeness, tone and policy compliance, with disagreements adjudicated. Distinguish cosmetic edits from substantive corrections. Track source support and appropriate abstention, not acceptance rate alone.

Required acceptance tests include:

  • A refund-request draft contains no commitment.
  • No model or pipeline component can invoke sending.
  • Unsupported facts cause omission or escalation.
  • Ticket updates invalidate stale drafts.
  • Human rerouting is not overwritten.
  • Provider failure leaves tickets accessible.
  • Duplicate events create neither duplicate drafts nor duplicate sends.
  • Agent permissions also constrain retrieved context.

Production expansion requires meeting quality thresholds and passing all safety-critical tests. Any unauthorized financial promise in evaluation blocks launch pending remediation.

8. Rollout, ownership and operations

Indicative eight-week plan:

  • Weeks 1–2: Support Operations approves taxonomy, saved replies, triage rules and baseline measurements; ML audits data; Engineering completes integration and security design.
  • Weeks 3–4: Build the pipeline, UI, evaluation harness and operational controls.
  • Weeks 5–6: Run offline evaluation and shadow mode without changing queues or showing drafts.
  • Weeks 7–8: Pilot with selected agents, then expand by queue only after gates pass.

Use a randomized holdout where practical, balanced by queue and shift, to compare handling time, response time, corrections and customer outcomes.

Engineering owns availability, access controls and rollback. ML owns evaluation, calibration and drift monitoring. Support Operations owns policy content, QA, agent training and staffing. Legal approves refund wording and provider terms.

Provide separate kill switches for routing and drafting. Suspected unauthorized commitments, data exposure or account-security mishandling trigger immediate suspension of the affected capability and incident review. Review corrections and rejected drafts weekly; retraining requires fresh holdout evaluation, not automatic ingestion of agent edits.

Claims the judge couldn’t find in the brief

  • The existing send action requires an explicit authenticated agent action.
  • There is an existing manual-triage destination, provisionally General.
  • There is an existing authorized refund process agents must use.

Grades and run details

Decision model 85 · LLM judge 7 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly7%
  • passAddresses the actual decision90%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty40%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable68%
  • passSpecifies AI failure behaviour100%
  • passDefines evals and a launch bar100%
  • partialSuccess settles trade-offs27%
  • partialProposes tests that could fail70%
Run
Run
#1
API response time
49 s
Submitted
29 Sept 2026

Needs more than a quick edit4 of 10 checks passed

Contains an unsupported current routing accuracy figure, missing evaluation plan/launch bar, and no trade‑off rules; a PM would need to re‑check evidence and add substantial missing sections.

Output

Product Requirements Document (PRD): AI-Powered Ticket Triage and Response Draft Generation ("AutoDraft & Route")

---

1. Document Overview

MetadataDetails
:---:---
Project NameAutoDraft & Route (Ledgerly Support Intelligence)
Target AudienceEngineering, Machine Learning (ML), and Support Operations (Ops) Leads
Document StatusDraft / Ready for Engineering Review
Target LaunchQ1 2027
GoalReduce Average First Response Time (FRT) from 7 hours to <2 hours while maintaining human-in-the-loop safety and strict legal compliance on financial commitments.

---

2. Background & Problem Statement

2.1 Current State & Metrics

Ledgerly is a premier accounting software platform designed specifically for small businesses. As our customer base scales, our support infrastructure is facing severe volume pressures: * Support Team Size: 42 agents distributed across four specialized queues: Billing, Technical, Account Access, and General. * Ticket Volume: ~9,000 incoming tickets per week, with 38% (~3,420 tickets/week) falling into the Billing queue. * First Response Time (FRT): Current average is 7 hours, severely lagging behind our target of under 2 hours. * Repetitive Workflows: Approximately 60% of all incoming tickets are resolved using one of roughly 40 standard saved replies (e.g., password resets, invoice retrieval, subscription tier changes).

2.2 The Opportunity

We possess two years of historical resolved tickets, including full conversation histories and agent replies. This corpus serves as a robust training and evaluation dataset for an ML-driven system that can: 1. Accurately classify and route incoming tickets to the correct queue. 2. Automatically draft context-aware first responses using our proven saved replies and historical resolutions.

2.3 Core Constraint (Legal Compliance)

  • Zero Automated Sending: Legal has mandated a strict human-in-the-loop requirement. The AI system must never autonomously send messages to customers.
  • Financial Safety: The model is strictly prohibited from making automated commitments regarding refunds, credits, or billing adjustments without explicit human review and approval.

---

3. Goals and Success Metrics (KPIs)

MetricCurrent BaselineTarget (3 Months Post-Launch)
:---:---:---
Average First Response Time (FRT)7 hours< 2 hours
Routing Accuracy~75% (Manual/Rule-based)1> 95% across all four queues
Draft Acceptance RateN/A> 70% (used as-is or with minor edits)
Agent Handle Time (AHT)Baseline-25% reduction in time spent crafting initial replies
Compliance / Unauthorized Sends00 violations (Hard safety guardrail)

---

4. User Personas & Workflows

4.1 Support Agent

  • Role: Resolves customer queries, reviews AI drafts, edits when necessary, approves responses, and handles complex escalations.
  • Workflow: Opens the ticketing dashboard, sees incoming tickets pre-sorted into their designated queue with an AI-generated draft response attached. The agent reviews the draft, clicks "Approve & Send" or modifies the text, ensuring zero financial compliance risk.

4.2 Support Operations Lead

  • Role: Manages queue health, tracks SLAs, audits agent performance, and maintains the library of saved replies and routing rules.
  • Workflow: Monitors routing precision dashboards, updates macro/saved reply templates, and analyzes draft acceptance/rejection feedback loops to refine prompts.

4.3 Engineering & ML Leads

  • Role: Build, deploy, scale, and monitor the inference pipeline, classification models, LLM orchestration, and UI integration.

---

5. Functional Requirements

5.1 Automated Ticket Routing (Triage Engine)

  • Queue Classification: Upon ticket ingestion, the ML classification service must categorize incoming text into one of four queues: `Billing`, `Technical`, `Account Access`, or `General`.
  • Confidence Thresholds:
  • Confidence $\ge$ 85%: Automatically assign the ticket to the predicted queue.
  • Confidence < 85%: Route the ticket to a "Triage Review" sub-queue or default General queue with a visual tag flagging low classification confidence.
  • Metadata Tagging: Tag tickets with sub-intent labels (e.g., `Billing: Refund Request`, `Technical: Bank Sync Failure`) to assist the response generation engine.

5.2 Response Draft Generation

  • Context Window & Retrieval: The generation engine must leverage the incoming user query, customer account metadata (plan tier, active status), and relevant historical examples from the 2-year resolved ticket dataset (via RAG / vector search).
  • Saved Reply Integration: The model must prioritize mapping inquiries to Ledgerly's 40 core saved replies where applicable, maintaining brand voice, accuracy, and tone consistency.
  • Draft Status: Every generated response must be marked with a distinct internal status: `Draft - Pending Agent Review`. It must remain locked from customer view until explicit human sign-off.

5.3 Agent Workspace UI/UX Integration

  • Side-by-Side Review: The agent UI must display the incoming customer message alongside the AI-generated draft in a clear, editable text box.
  • Action Controls:
  • [Approve & Send]: Immediately transmits the draft to the customer and logs the interaction.
  • [Edit & Send]: Allows agents to modify text inline before sending. Edits must be logged for ML fine-tuning feedback loops.
  • [Discard & Rewrite]: Clears the draft if the AI misunderstood the prompt.
  • Visual Safety Badges: Clear UI warnings must appear on any ticket flagged as containing billing or refund keywords, reminding agents of compliance protocols.

---

6. Non-Functional & Safety Requirements

6.1 Legal & Financial Guardrails (Critical)

  • No Auto-Dispatch: Zero API endpoints or automation rules are permitted to transition a draft state to `Sent` without a cryptographic token or database action originating from an authenticated human agent session.
  • Refund Keyword Interceptor: A deterministic rule-based filter must scan all generated drafts for high-risk financial terminology (`refund`, `credit`, `waiver`, `chargeback`). If detected, the UI must highlight the section in amber and require a secondary manual checkbox confirmation from the agent before sending.

6.2 Performance & Latency

  • Inference Latency: End-to-end classification and draft generation must complete within < 3.0 seconds of ticket ingestion via webhook, ensuring drafts are ready before an agent opens the ticket.
  • Availability: 99.9% uptime SLA for the ML inference and routing service.

6.3 Data Privacy & Security

  • PII Redaction: Customer tax IDs, banking credentials, and sensitive financial figures must be handled in compliance with GDPR, CCPA, and Ledgerly security policies. Training data pipelines must scrub extraneous PII prior to embedding generation.

---

7. Data, ML, & Training Strategy

7.1 Dataset Utilization

  • Corpus: 2 years of historical resolved tickets (approx. 900,000+ interaction pairs across 42 agents).
  • Data Splitting: 80% training/fine-tuning, 10% validation, 10% held-out test set reflecting recent seasonal billing spikes.

7.2 Model Architecture & Approach

  • Classification: Fine-tuned encoder model (e.g., BERT/DeBERTa variant) for multi-class queue routing and intent tagging.
  • Generation: Instruction-tuned LLM accessed via secure API or private VPC deployment, augmented with a vector database (RAG) containing the 40 approved saved replies and top-performing historical resolutions.
  • Feedback Loop: Log all agent edits (diffs) and rejections to construct a weekly RLHF (Reinforcement Learning from Human Feedback) or supervised fine-tuning dataset to reduce hallucination rates.

---

8. High-Level Architecture & Data Flow

``` [Customer Email / Web Form] │ ▼ [Ingestion Service / Webhook] │ ├──► [ML Classification Engine] ──► Assigns Queue & Sub-intent │ └──► [RAG Vector DB (Saved Replies & History)] │ ▼ [LLM Draft Generation Service] ──► Generates Initial Response Text │ ▼ (Strictly Locked: Status = Pending Review) [Support Ticketing Database] │ ▼ [Agent UI Dashboard] ◄─── Human Agent Reviews, Edits, & Approves │ ▼ (Explicit Human Action) [Customer Dispatch API] ```

---

9. Phased Rollout Plan

  • Phase 1: Shadow Mode (Weeks 1–4)
  • Deploy classification and draft generation in the background.
  • Log routing accuracy and draft quality without surfacing drafts to agents. Establish baseline ML accuracy metrics.
  • Phase 2: Internal Pilot / Beta (Weeks 5–8)
  • Roll out the UI feature to a pilot group of 5 senior Billing agents.
  • Measure acceptance rates, time saved, and friction points. Refine prompt engineering and guardrails.
  • Phase 3: General Availability Across All Queues (Weeks 9–12)2
  • Roll out to all 42 agents across Billing, Technical, Account Access, and General queues.
  • Establish continuous monitoring dashboards for Support Ops and Engineering leads.

---

10. Open Questions & Risks

  1. Edge Cases in Billing: How should the model handle complex multi-invoice dispute threads where historical context spans multiple months? (Mitigation: Surface the last 3 ticket summaries alongside the draft).
  2. Agent Adoption: Will agents trust the AI drafts, or will they rewrite them from scratch? (Mitigation: Emphasize time savings during training and incorporate agent feedback buttons).

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimStart again

    Source: Volumes9,000 tickets a week; 38% are billing.

    What we’d changeRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.

  2. 2Test or gate too weakTighten the testSubstantial rework

    What we’d changeAdvance each phase on quality gates (routing precision, draft acceptance, zero refund-commitment misses) rather than the calendar, with a human grading rubric and a rollback trigger.

Claims the judge couldn’t find in the brief

  • Current routing accuracy is ~75% (Manual/Rule-based).

Grades and run details

Decision model 55 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly63%
  • passAddresses the actual decision47%
  • partialRespects explicit constraints41%
  • partialIdentifies material uncertainty90%
  • partialAvoids unsupported claims48%
  • partialProduces the required deliverable43%
  • passSpecifies AI failure behaviour26%
  • partialDefines evals and a launch bar84%
  • partialSuccess settles trade-offs82%
  • partialProposes tests that could fail74%
Artefacts
Run
Run
#1
Time to output
14 s
Submitted
24 Sept 2026

Usable with a quick edit9 of 10 checks passed

The PRD is complete and actionable for the named leads, requiring only the addition of explicit numeric thresholds for already-defined gates. No evidence or reasoning would need to be redone.

Output

PRD: AI First-Response Drafting and Ticket Routing

Status: Proposed Product: Ledgerly Support Audience: Engineering, ML, and Support Operations leads

1. Summary

Ledgerly receives about 9,000 support tickets per week. Tickets are handled by 42 agents across Billing, Technical, Account access, and General. Average first response time is seven hours, against a target of under two hours. About 60% of tickets receive one of roughly 40 saved replies.

This feature will recommend a queue and prepare a first-response draft when a new ticket arrives. An agent must review and approve every draft before it is sent. The feature is intended to reduce time spent triaging and composing routine responses, while leaving decisions and customer communication under agent control. It must never send a response autonomously or make a refund commitment on Ledgerly’s behalf.

2. Goals and non-goals

Goals - Reduce median time from ticket creation to first response, with an operational goal of under two hours. - Reduce agent effort on routine first responses by using appropriate saved replies and ticket-specific details. - Route tickets to the best-fit queue, while making uncertain recommendations visible and easy to correct. - Preserve agent review and control over every customer-facing response.

Non-goals - Automatically send, resolve, or close tickets. - Decide refund eligibility, promise a refund, or commit to refund timing. - Draft replies to later messages in an existing conversation in the initial release. - Replace queue ownership, escalation policies, or agents’ judgment.

3. Users and workflow

The primary user is a support agent reviewing new tickets in their assigned queue. Support Operations owns queue definitions, saved replies, and handling guidance. ML and Engineering own model quality, serving, integrations, and monitoring.

For each new ticket, the system will: 1. Read the ticket’s permitted content and available support context. 2. Recommend one of the four queues, with a confidence indicator and brief reason. 3. Generate a first-response draft, preferentially based on a relevant approved saved reply where appropriate. 4. Display the recommendation and draft in the agent workflow. 5. Let the agent edit, discard, or approve. Approval sends the response through the existing support system; no response is sent before approval.1 6. Record the final queue, draft disposition, edits, and outcome for evaluation.

Agents may change the queue before or after reviewing the draft. A draft must remain clearly marked as AI-generated until approved.

4. Functional requirements

Routing

  • Classify each ticket as Billing, Technical, Account access, or General.
  • Show the recommended queue, confidence, and a short explanation grounded in ticket content.
  • Allow agents to override the recommendation; the override becomes an evaluation signal, not an automatic training label.
  • Route low-confidence or ambiguous cases to General and flag them for review. Support Operations must set and approve the launch confidence threshold using offline and pilot results.
  • Do not infer urgency or bypass existing escalation rules in the initial release.

Drafting

  • Generate a concise, relevant first response based on the ticket and approved support materials available to the system.
  • Use a saved reply when it fits, adapting only with verified ticket details. Do not fabricate account status, actions taken, policy, or resolution.
  • When information is insufficient, ask a clear clarifying question or provide a safe acknowledgement rather than guessing.
  • Support agents can edit, discard, or approve the draft. No draft may be sent without an explicit agent action.
  • If the request concerns a refund, the draft must not promise, confirm, or imply a refund or refund timing. It should use approved noncommittal wording and leave eligibility and timing to an agent. Refund-related tickets should be visibly flagged for agent attention.

5. Data, model, and system requirements

Ledgerly has two years of resolved tickets and agent replies. ML should assess data coverage and quality before training: queue labels, duplicate or reopened cases, outdated answers, agent-specific wording, and tickets whose resolution depended on account information unavailable at ticket creation. Historical agent replies are examples, not policy; approved saved replies and current support guidance take precedence.

Create a time-based training, validation, and test split to reduce leakage from repeated tickets or changing policies. Evaluate routing and draft quality separately, including by queue, ticket type, and relevant risk category. Do not train on post-response information when evaluating first-response behavior. Remove or protect unnecessary sensitive data, and use only the customer and account context needed for support.

The serving path must retrieve only context the agent is authorized to see. Engineering must confirm integration points with the ticketing system, queue assignment, saved replies, and send/approval controls before implementation. Log model/version, recommendation, confidence, draft, agent edits, approval or discard, final queue, and response timestamps, subject to Ledgerly’s retention and access policies. Restrict access to these logs and avoid storing secrets or unnecessary personal data in model prompts.

6. Quality, safety, and success measures

Establish a pre-launch baseline and report weekly: - Median and 90th-percentile time to first response; proportion answered within two hours. - Queue recommendation accuracy and agent override rate, overall and by queue. - Draft approval rate, edit rate, discard rate, and time from ticket opening to send. - Agent-rated usefulness and sampled quality review results. - Refund-policy violations, fabricated claims, and other serious customer-impacting errors.

Suggested launch gates, to be confirmed by Support Operations and Legal: routing performance must be no worse than the current human baseline on the held-out set; pilot draft samples must meet an agreed quality rubric; and no refund commitment may appear in an approved draft during safety testing. Any refund commitment, unauthorized send, or material fabricated claim triggers immediate disablement of the affected capability and incident review. Do not optimize approval rate alone: agents must be able to reject drafts without penalty.

7. Rollout and operations

  1. Discovery and offline evaluation: Validate data, current queue definitions, approved saved replies, policy ownership, and integration feasibility. Build a labeled test set reviewed by Support Operations, with explicit refund and other high-risk examples.
  2. Shadow mode: Run recommendations without showing or applying them. Compare predicted queues and drafts with agent outcomes; inspect errors and subgroup performance.
  3. Limited pilot: Enable the workflow for a small, representative group of agents and ticket types. Keep agent approval mandatory. Support Operations reviews a daily sample and gathers structured agent feedback.
  4. Expansion: Expand to all 42 agents only after launch gates are met and Legal and Support Operations approve. Roll back by disabling drafting or routing independently if either degrades.

Support Operations owns the queue taxonomy, approved reply content, quality rubric, and agent training. ML owns evaluation, thresholds, drift monitoring, and model updates. Engineering owns access controls, reliable delivery, audit logs, feature flags, and rollback. A named on-call owner must be assigned before pilot launch. Review quality and performance weekly during rollout and after material model, policy, or saved-reply changes. Re-evaluate against a fresh sample regularly; do not promote agent edits to training data without review and policy checks.

8. Open decisions before build

Confirm the ticketing platform and integration surfaces; which customer/account fields are available at ticket creation; the authoritative source for current support policy; Legal’s approved refund language and review process; the minimum acceptable routing and draft-quality thresholds; and the retention period for prompts, drafts, and audit logs. These decisions are launch prerequisites, not reasons to relax agent approval or refund safeguards.

What a PM had to fix

  1. 1Test or gate too weakTighten the testTargeted repair

    What we’d changeAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly23%
  • passAddresses the actual decision66%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims39%
  • passProduces the required deliverable37%
  • passSpecifies AI failure behaviour98%
  • passDefines evals and a launch bar45%
  • partialSuccess settles trade-offs52%
  • partialProposes tests that could fail77%
Run
Run
#1
API response time
30 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongWrongRight
GPT-6.1 Sol · API

It invents current systems not in the brief, such as an existing authenticated send action, manual-triage destination, and authorized refund process.

Gemini 3.5 Flash-Lite · Gemini

States current routing accuracy of ~75% as fact without any support from the supplied evidence and does not label it as an assumption.

GPT-6 Luna · API

All current-state figures used come directly from the supplied context, with no invented facts about the current situation.

Identifies material uncertaintyRightWrongRight
GPT-6.1 Sol · API

It names missing channels/languages, unmeasured thresholds, sample-size limitations, and how evaluation and pilot results would resolve them.

Gemini 3.5 Flash-Lite · Gemini

Open questions are listed but are not tied to decision‑changing thresholds or explicit resolution plans; the PRD does not state what would halt or alter the rollout.

GPT-6 Luna · API

It identifies material unknowns such as data quality, integration points, policy ownership, and confidence thresholds, and says how to resolve them through evaluation and piloting.

Avoids unsupported claimsRightWrongRight
GPT-6.1 Sol · API

Normative requirements and design rules are framed as proposals, not as established facts from the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

The routing accuracy baseline of ~75% is presented as a fact without being marked as an assumption, and no evidence supports it.

GPT-6 Luna · API

No causes, forecasts, or interpretations are presented as established fact; uncertain items are framed as suggested, open, or to be confirmed.

Defines evals and a launch barRightWrongRight
GPT-6.1 Sol · API

It defines an adjudicated test set, quality metrics, thresholds, and launch gates.

Gemini 3.5 Flash-Lite · Gemini

No evaluation dataset criteria, metrics (e.g., draft quality), or specific launch threshold for progressing from shadow mode to pilot or GA are defined.

GPT-6 Luna · API

It defines a pre-launch baseline, metrics, a held-out test set, launch gates, and a zero-refund-commitment safety condition.

Success settles trade-offsWrongWrongRight
GPT-6.1 Sol · API

It does not give an explicit trade-off rule such as accepting lower coverage to preserve precision above a stated level.

Gemini 3.5 Flash-Lite · Gemini

Success metrics include targets and a timeframe, but no explicit trade‑off rule (e.g., coverage vs. precision) is stated.

GPT-6 Luna · API

It names a success metric with an under-two-hours target and weekly cadence, and gives trade-off rules such as disabling the affected capability on any refund commitment or unauthorized send.

All got wrong 1

Proposes tests that could failWrongWrongWrong
GPT-6.1 Sol · API

Several acceptance tests and kill criteria lack numeric thresholds, measurement windows, or the action each outcome triggers.

Gemini 3.5 Flash-Lite · Gemini

The phased rollout lacks numeric thresholds, measurement windows, and actions tied to results; it does not define kill criteria.

GPT-6 Luna · API

Proposed gates lack numeric thresholds and explicit time windows; 'no worse than human baseline' and 'agreed quality rubric' are not measurable fail criteria, and pilot/shadow durations are unspecified.

All got right 4

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

It commits to an MVP with a clear approval flow, release gates, and conditions that would change rollout.

Gemini 3.5 Flash-Lite · Gemini

The PRD commits to a clear design for the AI feature, framed for the engineering, ML, and ops leads.

GPT-6 Luna · API

The PRD commits unambiguously to AI-assisted queue recommendation and draft generation with mandatory agent approval, and specifies conditions for rollback or disablement.

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

It respects the named audience, length, human approval requirement, and legal prohibition on automated refund commitments.

Gemini 3.5 Flash-Lite · Gemini

The document remains within the 1,000–1,500‑word range, addresses the required readers, and enforces legal constraints with human‑only dispatch and a refund‑keyword interceptor.

GPT-6 Luna · API

It respects the named audience, length, PRD form, and enforces the no-automated-refund-commitment constraint through mandatory approval, noncommittal refund wording, and flagging.

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

It is a complete PRD for engineering, ML, and support-operations leads that could start work from.

Gemini 3.5 Flash-Lite · Gemini

The output is a complete PRD with functional requirements, architecture, and rollout plan that the named leads could start work from.

GPT-6 Luna · API

The requested PRD is present, complete, written for engineering, ML, and support-operations leads, and within the requested length.

Specifies AI failure behaviourRightRightRight
GPT-6.1 Sol · API

It specifies low-confidence routing to manual triage, abstention or clarification when evidence is missing, stale-draft disabling, and outage fallback.

Gemini 3.5 Flash-Lite · Gemini

Low-confidence classification routes to a triage queue, and a deterministic refund‑keyword interceptor triggers a secondary manual checkbox.

GPT-6 Luna · API

The PRD specifies low-confidence routing to General with flags, clarifying questions when information is insufficient, and immediate disablement for serious errors.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT89.4100.02None
2Sonnet 5.5withAPI78.675.52None
3GPT-6.1 SolwithAPI78.661.82None
4GPT-6 LunawithAPI81.155.52None
5Opus 5.5withClaude73.661.42None
6Gemini 3.5 Flash-LitewithGemini52.538.22None
7Gemini 3.8 FlashwithAPI57.523.22None

About the task

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Defines success precisely enough to settle trade-offs
  • Says what is out of scope

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

AI product PRD (core) · Conventional product PRD