Failure patterns

These are the mistakes AI keeps making on real PM work, and what to do about each. Every one comes from a blind review where a PM recorded exactly what they had to fix: the passage, the source it contradicts, and the change.

Issue detail covers 73 of 135 published outputs.

Who makes which mistake

Every setup has its own tells. Here’s how many of each one’s reviewed outputs made each mistake at least once.

Outputs with each mistake, by setup
MistakeOpus 5.5ClaudeGPT-6 AstraChatGPTGPT-6.1 SolAPIGPT-6 LunaAPISonnet 5.5APIGemini 3.5 Flash-LiteGemini
Hypothesis stated as factReframe it as a hypothesis11 of 262 of 19none of 24none of 23none of 2413 of 19
Invented evidenceVerify or remove the claim7 of 261 of 19none of 243 of 23none of 2411 of 19
Numbers wrongRedo the arithmetic5 of 261 of 19none of 24none of 231 of 246 of 19
Test or gate too weakTighten the test5 of 264 of 19none of 242 of 23none of 242 of 19
Contradiction missedSurface the contradiction1 of 263 of 19none of 241 of 23none of 244 of 19
Constraint missedRestore the constraint4 of 26none of 19none of 24none of 23none of 244 of 19
Decision deferredMake the callnone of 265 of 19none of 24none of 23none of 24none of 19
Artefact unusableFix the artefactnone of 261 of 19none of 24none of 23none of 242 of 19
Dot area is the share of that setup’s reviewed outputs with the mistake at least once.

Hypothesis stated as fact

Presents a plausible explanation or cause as if the evidence had established it.

What you doReframe it as a hypothesis

Gemini 3.5 Flash-LitewithGemini on Supply or demand for a stalled marketplace
It wrote
The failure is entirely in initial activation and discovery.
The source (Research)

Students cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.

Reframe it as a hypothesis · start againPresent matching friction as the hypothesis the research suggests, not a proven diagnosis, and stage the spending so the board learns before it commits the full £1.2m.

Read the output
Opus 5.5withClaude on AI triage for support tickets
It wrote
Most of that delay is not writing time. It is triage and lookup
The source (Volumes)

Average first response is 7 hours against a target of under 2.

Reframe it as a hypothesis · substantial reworkPresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.

Read the output
GPT-6 AstrawithChatGPT on The CEO's embedded-payments bet
It wrote
the supplied evidence supports an illustrative annual revenue range of $3.3M–$7.2M
What was missing

Presents a plausible explanation or cause as if the evidence had established it.

Reframe it as a hypothesis · quick editFrame the range up front as illustrative full-capture scenarios, with actual capture still unmeasured, so it isn't read as a forecast.

Read the output

Invented evidence

States a fact, figure, quote or current behaviour the source material doesn’t contain.

What you doVerify or remove the claim

Gemini 3.5 Flash-LitewithGemini on AI triage for support tickets
It wrote
~75% (Manual/Rule-based)
The source (Volumes)

9,000 tickets a week; 38% are billing.

Verify or remove the claim · start againRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.

Read the output
Opus 5.5withClaude on Go/no-go for AI-drafted support replies
It wrote
Training is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.
The source (Support operations)

Training for all 42 agents is booked for Friday.

Verify or remove the claim · targeted repairThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.

Read the output
GPT-6 LunawithAPI on Supply or demand for a stalled marketplace
It wrote
Allocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.
What was missing

States a fact, figure, quote or current behaviour the source material doesn’t contain.

Verify or remove the claim · targeted repairTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.

Read the output

Numbers wrong

Misreads the data, mixes up steps or denominators, or gets the sums wrong.

What you doRedo the arithmetic

Gemini 3.5 Flash-LitewithGemini on Fitness app first week
It wrote
while 52% start the trial, only 19% complete three workouts in week one
The source (Funnel)

started trial 52% → first workout 48% → three workouts in week one 19%

Redo the arithmetic · start againRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.

Read the output
Opus 5.5withClaude on Eight calls with finance teams
It wrote
Six of eight say reconciliation is a significant part of close (Calls 1, 3, 4, 5, 7, 8).
The source (Call 4: CFO, manufacturing company (250 staff))

We switched reconciliation tools last year and I regret it.

Redo the arithmetic · targeted repairRecount without Call 4, which is about a failed tool switch, not close time. It's five of eight.

Read the output
GPT-6 AstrawithChatGPT on The underpowered onboarding test
It wrote
Extend the experiment instead if a possible 1.4pp activation loss is commercially unacceptable
The source (Readout)

Activation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

Redo the arithmetic · quick editDescribe 1.4pp as the lower end of the interval, not the largest possible loss, and qualify the one-day cost assumption.

Read the output

Test or gate too weak

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

What you doTighten the test

Gemini 3.5 Flash-LitewithGemini on An AI SDR for small agencies
It wrote
If agencies will not engage with or pre-order a warm-relationship retention tool at $200–$600/month, they will certainly reject a cold AI prospecting tool.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · substantial reworkTest the reason the memo gives: whether agencies without a BD person can turn AI-sent cold email into deals. Demand for a different product doesn't settle that.

Read the output
GPT-6 LunawithAPI on AI triage for support tickets
It wrote
Approval sends the response through the existing support system; no response is sent before approval.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · targeted repairAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.

Read the output
GPT-6 AstrawithChatGPT on AI triage for support tickets
It wrote
at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · quick editSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.

Read the output

Constraint missed

Breaks or ignores something the brief required or ruled out, or adds scope nobody asked for.

What you doRestore the constraint

Gemini 3.5 Flash-LitewithGemini on Meeting summaries with action items
It wrote
Ground-truth list of officially invited participants (Email, Name, Role).
The source (Constraint)

Summaries must never assign an action to someone who was not in the meeting.

Restore the constraint · start againBuild the list of possible owners from who actually joined, not who was invited, and enforce it in code on generated owners, manual edits and sending. A model-set is_verified_attendee flag can't guarantee it.

Read the output
Opus 5.5withClaude on Meeting summaries with action items
It wrote
The roster filter runs as deterministic code after the model, never inside the prompt alone.
The source (Constraint)

Summaries must never assign an action to someone who was not in the meeting.

Restore the constraint · targeted repairApply the roster check to organiser edits and the final send as well, not only to the model's output.

Read the output
Gemini 3.5 Flash-LitewithGemini on Expense receipt capture
It wrote
You can preview the interactive dashboard
The source (Brief)

captures a receipt, checks the amount the app extracted, corrects it, and saves the expense.

Restore the constraint · quick editCut the dashboard and the finance-approval step, and the confusing Cancel after saving. The brief asks only to capture, correct and save.

Read the output

Contradiction missed

Smooths over conflicting evidence that changes the answer.

What you doSurface the contradiction

Gemini 3.5 Flash-LitewithGemini on Conversion up, retention down
It wrote
Variant B’s retention dropped by 4.5pp
The source (Readout summary)

Test ran 21 days

Surface the contradiction · substantial reworkFlag that a 21-day test can't have complete day-30 outcomes, and ask when the readout was produced before leaning on the retention figure.

Read the output
GPT-6 AstrawithChatGPT on Fitness app first week
It wrote
We also lose 29% before onboarding finishes.
The source (Funnel)

Install → finished onboarding 71% → started trial 52%

Surface the contradiction · targeted repairAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.

Read the output
Opus 5.5withClaude on Conversion up, retention down
It wrote
Day-30 retention: −4.5pp against a 2pp limit.
The source (Readout summary)

Test ran 21 days

Surface the contradiction · quick editAsk how a 21-day test reports day-30 retention before relying on the figure.

Read the output

Decision deferred

Lays out options or hedges where the brief asked for a decision the evidence can support.

What you doMake the call

GPT-6 AstrawithChatGPT on Go/no-go for AI substitutions, from the rollout data
It wrote
Expand only where the evidence supports it; insufficient evidence means continued limited exposure.
What was missing

Lays out options or hedges where the brief asked for a decision the evidence can support.

Make the call · targeted repairGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.

Read the output
GPT-6 AstrawithChatGPT on Conversion up, retention down
It wrote
Its confidence interval includes declines within our tolerance, so we cannot say the true loss definitively exceeds 2pp.
The source (Guardrails agreed before launch)

Day-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.

Make the call · targeted repairMake the argument rather than hedging it: the point estimate breaches the guardrail and most of the interval is past it. The refund question can be settled from the numbers given, so settle it.

Read the output
GPT-6 AstrawithChatGPT on The CEO asks why a metric dropped
It wrote
Navigation may be contributing, but we don’t yet know how much.
The source (Metrics)

Mobile only: −5.2%.

Make the call · quick editAnswer Lena's question more directly: mobile's 5.2% drop with no navigation change points away from the navigation as the main cause. Tighten the SSO hedge.

Read the output

Other

Anything else a PM had to fix.

What you doFix it

GPT-6 LunawithAPI on The CEO's embedded-payments bet
It wrote
Together, that supports roughly $4.5M–$10M of annualized revenue, before accounting for contract lock-in or slower adoption among large accounts.
What was missing

Anything else a PM had to fix.

Fix it · targeted repairThese assumptions produce an illustrative $4.5M–$10M envelope, not a validated revenue range. They assume customer adoption translates proportionally into invoice volume and the modelled payments route through Pay. Rebuild by segment: accessible invoice value × volume-weighted adoption × routing share × payment-method economics, excluding contracted volume.

Read the output
Opus 5.5withClaude on Analytics tool losing users at setup
It wrote
Instrument every screen, 1 to 11, before shipping anything.
What was missing

Anything else a PM had to fix.

Fix it · quick editInstrument alongside fix 1 rather than before it, and merge fixes 1 and 4, which both land the user on an auto-drawn chart.

Read the output
Gemini 3.5 Flash-LitewithGemini on The underpowered onboarding test
It wrote
Continuing to run the test for another four weeks just to chase statistical significance on a small effect isn't a wise use of time
The source (Options on the table)

Running the test for another four weeks at current traffic would detect an effect of about 2pp.

Fix it · targeted repairWeigh the risk of a small activation loss, not only the one day of engineering, and say what result would justify the extra four weeks.

Read the output

Artefact unusable

The prototype or deck doesn’t work, breaks the format asked for, or couldn’t be shared as delivered.

What you doFix the artefact

GPT-6 AstrawithChatGPT on Clinic appointment rebooking
It wrote
Includes seven days of availability, fully booked states
The source (Constraints)

Show a clear "no slots available" state for a fully booked day.

Fix the artefact · quick editMark fully booked days in the day picker, so patients don't have to tap each day to find out.

Read the output
Gemini 3.5 Flash-LitewithGemini on Clinic appointment rebooking
It wrote
with a single-tap jump to the next available slot.
The source (Constraints)

Show a clear "no slots available" state for a fully booked day.

Fix the artefact · quick editFix the jump from a fully booked day: it suggests the wrong next available date. Everything else is ready to test.

Read the output
Gemini 3.5 Flash-LitewithGemini on Expense receipt capture
It wrote
observe the AI misread value of £83.40 with the alert badge
What was missing

The prototype or deck doesn’t work, breaks the format asked for, or couldn’t be shared as delivered.

Fix the artefact · quick editWord the warning as 'Please check this amount' rather than revealing the correct figure, or the review step tests nothing.

Read the output