Claude Opus 5.5
Anthropic · claude-opus-5-5 · effort medium (default)
What you can hand it, and what you’ll still have to catch
Dropped 27 Sept 2026, against whatever leads today.
Makes clear calls, but check the evidence behind them.
Opus 5.5 commits to a recommendation in every case that asks for one, and its best memos rebuild the numbers from scratch. Seven of its 18 outputs were usable with at most a quick edit. The usual fix is a targeted repair: it states its explanations as findings, and fills gaps in the brief with facts of its own.
What you can hand it
- One-shot prototypes. Both were rated Ready to use.
- Challenging a plan with the numbers. It found that the payments model applied a card-only margin to every invoice, and re-estimated the revenue from the data.
- Launch calls from raw data. It worked the rollout figures region by region and answered each stakeholder's claim with numbers that check out.
- Memos that have to choose. It picked a side in every strategy, launch and experiment case.
What you’ll still have to catch
- Explanations presented as findings. We recorded this in 11 of its 18 outputs, for example "The problem is matching, not demand".
- Facts the brief never gave: systems that "already exist", data "we don't track today", customers who "churn or ask for refunds".
- Statistical overreach: a confidence interval read as ruling out any effect, and "the nav can explain at most about 1 point".
- Dates and schedules it fills in itself, then plans around.
- Best setupOpus 5.5 · Claude73.2#1 overall
Task by task
Its Combined score on each core task, across 26 outputs. Beside it, the mistakes a PM recorded in earlier reviews.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
- Hypothesis stated as fact11 of 26Develop product strategy, Extract discovery insights, Write a PRD
- Invented evidence7 of 26Extract discovery insights, Develop product strategy, Write a stakeholder update
- Test or gate too weak5 of 26Make the launch call, Challenge an idea, Develop product strategy
- Numbers wrong5 of 26Analyse experiment results, Write a stakeholder update, Extract discovery insights
- Constraint missed4 of 26Write a PRD, Make the launch call, Extract discovery insights
- Other2 of 26Challenge an idea, Activation & onboarding review
- Contradiction missed1 of 26Analyse experiment results
What you’ll need to check
The fixes you’ll make most often, each with a real example.
Reframe it as a hypothesis
Presents a plausible explanation or cause as if the evidence had established it.
Most of that delay is not writing time. It is triage and lookup
Average first response is 7 hours against a target of under 2.
Reframe it as a hypothesis · substantial reworkPresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.
Read the outputVerify or remove the claim
States a fact, figure, quote or current behaviour the source material doesn’t contain.
Training is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.
Training for all 42 agents is booked for Friday.
Verify or remove the claim · targeted repairThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.
Read the outputTighten the test
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
as does a treatment refund rate more than 10% above its holdout over a rolling 3 days.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · quick editSet the refund trigger outside day-to-day noise, with a bigger gap or a longer window, or it will switch regions off for no reason.
Read the outputHow it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap72.5
- Customer research call guide85.6
- Experiment specification76.9
- Develop product strategy73.9
- One-shot prototype76.8
- Write a stakeholder update64.4
- Make the launch call84.2
- 1000x an idea82.0
- Write a PRD64.0
- Extract discovery insights71.3
- Challenge an idea63.8
- Analyse experiment results63.1
- Activation & onboarding review73.8
Opus 5.5 · Claude Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| 1 | Opus 5.5withClaudeClaude 2026-09 | 82.1 | 70.1 |
Strengths
By category
Opus 5.5 · Claude
Consistency
Typical time to output
Published runs
26 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| 1000x an ideaFrom tip calculator to worker network | Claude | |
| 1000x an ideaSharing a shopping list, 1000x | Claude | |
| Activation & onboarding reviewAnalytics tool losing users at setup | Claude | |
| Activation & onboarding reviewFitness app first week | Claude | |
| Analyse experiment resultsConversion up, retention down | Claude | |
| Analyse experiment resultsThe underpowered onboarding test | Claude | |
| Build a roadmapA year of spend management, with a hard deadline | Claude | |
| Build a roadmapTwo squads, eight asks, one half | Claude | |
| Challenge an ideaAn AI SDR for small agencies | Claude | |
| Challenge an ideaThe CEO's embedded-payments bet | Claude | |
| Customer research call guideWhy are warehouses leaving? Two execs, two theories | Claude | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | Claude | |
| Develop product strategyHidden validation case | Claude | |
| Develop product strategySupply or demand for a stalled marketplace | Claude | |
| Experiment specificationShowing the delivery fee up front | Claude | |
| Experiment specificationBatching deliveries before peak season | Claude | |
| Extract discovery insightsEight calls with finance teams | Claude | |
| Extract discovery insightsHidden validation case | Claude | |
| Make the launch callGo/no-go for AI-drafted support replies | Claude | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | Claude | |
| One-shot prototypeExpense receipt capture | Claude | |
| One-shot prototypeClinic appointment rebooking | Claude | |
| Write a PRDMeeting summaries with action items | Claude | |
| Write a PRDAI triage for support tickets | Claude | |
| Write a stakeholder updateA quarterly business review deck, in brand | Claude | |
| Write a stakeholder updateThe CEO asks why a metric dropped | Claude |