Failure patterns
These are the mistakes AI keeps making on real PM work, and what to do about each. Every one comes from a blind review where a PM recorded exactly what they had to fix: the passage, the source it contradicts, and the change.
Issue detail covers 73 of 135 published outputs.
Who makes which mistake
Every setup has its own tells. Here’s how many of each one’s reviewed outputs made each mistake at least once.
| Mistake | Opus 5.5Claude | GPT-6 AstraChatGPT | GPT-6.1 SolAPI | GPT-6 LunaAPI | Sonnet 5.5API | Gemini 3.5 Flash-LiteGemini |
|---|---|---|---|---|---|---|
| Hypothesis stated as factReframe it as a hypothesis | 11 of 26 | 2 of 19 | none of 24 | none of 23 | none of 24 | 13 of 19 |
| Invented evidenceVerify or remove the claim | 7 of 26 | 1 of 19 | none of 24 | 3 of 23 | none of 24 | 11 of 19 |
| Numbers wrongRedo the arithmetic | 5 of 26 | 1 of 19 | none of 24 | none of 23 | 1 of 24 | 6 of 19 |
| Test or gate too weakTighten the test | 5 of 26 | 4 of 19 | none of 24 | 2 of 23 | none of 24 | 2 of 19 |
| Contradiction missedSurface the contradiction | 1 of 26 | 3 of 19 | none of 24 | 1 of 23 | none of 24 | 4 of 19 |
| Constraint missedRestore the constraint | 4 of 26 | none of 19 | none of 24 | none of 23 | none of 24 | 4 of 19 |
| Decision deferredMake the call | none of 26 | 5 of 19 | none of 24 | none of 23 | none of 24 | none of 19 |
| Artefact unusableFix the artefact | none of 26 | 1 of 19 | none of 24 | none of 23 | none of 24 | 2 of 19 |
Hypothesis stated as fact
Presents a plausible explanation or cause as if the evidence had established it.
What you doReframe it as a hypothesis
By setup
By task
- Extract discovery insights9 in 12
- Develop product strategy8 in 12
- Write a stakeholder update4 in 12
- Challenge an idea4 in 12
- Analyse experiment results4 in 12
- Activation & onboarding review4 in 12
- Make the launch call3 in 12
- Write a PRD3 in 12
The failure is entirely in initial activation and discovery.
Students cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Reframe it as a hypothesis · start againPresent matching friction as the hypothesis the research suggests, not a proven diagnosis, and stage the spending so the board learns before it commits the full £1.2m.
Read the outputMost of that delay is not writing time. It is triage and lookup
Average first response is 7 hours against a target of under 2.
Reframe it as a hypothesis · substantial reworkPresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.
Read the outputthe supplied evidence supports an illustrative annual revenue range of $3.3M–$7.2M
Presents a plausible explanation or cause as if the evidence had established it.
Reframe it as a hypothesis · quick editFrame the range up front as illustrative full-capture scenarios, with actual capture still unmeasured, so it isn't read as a forecast.
Read the outputInvented evidence
States a fact, figure, quote or current behaviour the source material doesn’t contain.
What you doVerify or remove the claim
By setup
By task
- Write a stakeholder update5 in 12
- Extract discovery insights5 in 12
- Develop product strategy4 in 12
- Make the launch call3 in 12
- Write a PRD3 in 12
- Activation & onboarding review3 in 12
- Challenge an idea2 in 12
- Analyse experiment results1 in 12
~75% (Manual/Rule-based)
9,000 tickets a week; 38% are billing.
Verify or remove the claim · start againRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.
Read the outputTraining is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.
Training for all 42 agents is booked for Friday.
Verify or remove the claim · targeted repairThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.
Read the outputAllocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.
States a fact, figure, quote or current behaviour the source material doesn’t contain.
Verify or remove the claim · targeted repairTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.
Read the outputNumbers wrong
Misreads the data, mixes up steps or denominators, or gets the sums wrong.
What you doRedo the arithmetic
By setup
By task
- Analyse experiment results4 in 12
- Challenge an idea3 in 12
- Extract discovery insights2 in 12
- Activation & onboarding review2 in 12
- Develop product strategy1 in 12
- Write a stakeholder update1 in 12
- Make the launch call1 in 12
while 52% start the trial, only 19% complete three workouts in week one
started trial 52% → first workout 48% → three workouts in week one 19%
Redo the arithmetic · start againRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.
Read the outputSix of eight say reconciliation is a significant part of close (Calls 1, 3, 4, 5, 7, 8).
We switched reconciliation tools last year and I regret it.
Redo the arithmetic · targeted repairRecount without Call 4, which is about a failed tool switch, not close time. It's five of eight.
Read the outputExtend the experiment instead if a possible 1.4pp activation loss is commercially unacceptable
Activation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).
Redo the arithmetic · quick editDescribe 1.4pp as the lower end of the interval, not the largest possible loss, and qualify the one-day cost assumption.
Read the outputTest or gate too weak
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
What you doTighten the test
By setup
By task
- Challenge an idea4 in 12
- Make the launch call3 in 12
- Write a PRD3 in 12
- Develop product strategy2 in 12
- Extract discovery insights1 in 12
- Activation & onboarding review1 in 12
If agencies will not engage with or pre-order a warm-relationship retention tool at $200–$600/month, they will certainly reject a cold AI prospecting tool.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · substantial reworkTest the reason the memo gives: whether agencies without a BD person can turn AI-sent cold email into deals. Demand for a different product doesn't settle that.
Read the outputApproval sends the response through the existing support system; no response is sent before approval.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · targeted repairAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.
Read the outputat least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · quick editSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.
Read the outputConstraint missed
Breaks or ignores something the brief required or ruled out, or adds scope nobody asked for.
What you doRestore the constraint
By setup
By task
- Write a PRD5 in 12
- Make the launch call3 in 12
- Develop product strategy1 in 12
- One-shot prototype1 in 12
- Extract discovery insights1 in 12
Ground-truth list of officially invited participants (Email, Name, Role).
Summaries must never assign an action to someone who was not in the meeting.
Restore the constraint · start againBuild the list of possible owners from who actually joined, not who was invited, and enforce it in code on generated owners, manual edits and sending. A model-set is_verified_attendee flag can't guarantee it.
Read the outputThe roster filter runs as deterministic code after the model, never inside the prompt alone.
Summaries must never assign an action to someone who was not in the meeting.
Restore the constraint · targeted repairApply the roster check to organiser edits and the final send as well, not only to the model's output.
Read the outputYou can preview the interactive dashboard
captures a receipt, checks the amount the app extracted, corrects it, and saves the expense.
Restore the constraint · quick editCut the dashboard and the finance-approval step, and the confusing Cancel after saving. The brief asks only to capture, correct and save.
Read the outputContradiction missed
Smooths over conflicting evidence that changes the answer.
What you doSurface the contradiction
By setup
By task
- Analyse experiment results3 in 12
- Develop product strategy2 in 12
- Make the launch call2 in 12
- Challenge an idea1 in 12
- Activation & onboarding review1 in 12
Variant B’s retention dropped by 4.5pp
Test ran 21 days
Surface the contradiction · substantial reworkFlag that a 21-day test can't have complete day-30 outcomes, and ask when the readout was produced before leaning on the retention figure.
Read the outputWe also lose 29% before onboarding finishes.
Install → finished onboarding 71% → started trial 52%
Surface the contradiction · targeted repairAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.
Read the outputDay-30 retention: −4.5pp against a 2pp limit.
Test ran 21 days
Surface the contradiction · quick editAsk how a 21-day test reports day-30 retention before relying on the figure.
Read the outputDecision deferred
Lays out options or hedges where the brief asked for a decision the evidence can support.
What you doMake the call
By setup
By task
- Write a stakeholder update2 in 12
- Make the launch call1 in 12
- Analyse experiment results1 in 12
- Activation & onboarding review1 in 12
Expand only where the evidence supports it; insufficient evidence means continued limited exposure.
Lays out options or hedges where the brief asked for a decision the evidence can support.
Make the call · targeted repairGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.
Read the outputIts confidence interval includes declines within our tolerance, so we cannot say the true loss definitively exceeds 2pp.
Day-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
Make the call · targeted repairMake the argument rather than hedging it: the point estimate breaches the guardrail and most of the interval is past it. The refund question can be settled from the numbers given, so settle it.
Read the outputNavigation may be contributing, but we don’t yet know how much.
Mobile only: −5.2%.
Make the call · quick editAnswer Lena's question more directly: mobile's 5.2% drop with no navigation change points away from the navigation as the main cause. Tighten the SSO hedge.
Read the outputOther
Anything else a PM had to fix.
What you doFix it
By setup
By task
- Challenge an idea2 in 12
- Make the launch call1 in 12
- Analyse experiment results1 in 12
- Activation & onboarding review1 in 12
Together, that supports roughly $4.5M–$10M of annualized revenue, before accounting for contract lock-in or slower adoption among large accounts.
Anything else a PM had to fix.
Fix it · targeted repairThese assumptions produce an illustrative $4.5M–$10M envelope, not a validated revenue range. They assume customer adoption translates proportionally into invoice volume and the modelled payments route through Pay. Rebuild by segment: accessible invoice value × volume-weighted adoption × routing share × payment-method economics, excluding contracted volume.
Read the outputInstrument every screen, 1 to 11, before shipping anything.
Anything else a PM had to fix.
Fix it · quick editInstrument alongside fix 1 rather than before it, and merge fixes 1 and 4, which both land the user on an auto-drawn chart.
Read the outputContinuing to run the test for another four weeks just to chase statistical significance on a small effect isn't a wise use of time
Running the test for another four weeks at current traffic would detect an effect of about 2pp.
Fix it · targeted repairWeigh the risk of a small activation loss, not only the one day of engineering, and say what result would justify the extra four weeks.
Read the outputArtefact unusable
The prototype or deck doesn’t work, breaks the format asked for, or couldn’t be shared as delivered.
What you doFix the artefact
By setup
By task
- One-shot prototype3 in 12
Includes seven days of availability, fully booked states
Show a clear "no slots available" state for a fully booked day.
Fix the artefact · quick editMark fully booked days in the day picker, so patients don't have to tap each day to find out.
Read the outputwith a single-tap jump to the next available slot.
Show a clear "no slots available" state for a fully booked day.
Fix the artefact · quick editFix the jump from a fully booked day: it suggests the wrong next available date. Everything else is ready to test.
Read the outputobserve the AI misread value of £83.40 with the alert badge
The prototype or deck doesn’t work, breaks the format asked for, or couldn’t be shared as delivered.
Fix the artefact · quick editWord the warning as 'Please check this amount' rather than revealing the correct figure, or the review step tests nothing.
Read the output