Ratings / Anthropic

Claude Opus 5.5

Anthropic · claude-opus-5-5 · effort medium (default)

Opus 5.5withClaudeRanked #1 overall
26 runs · ±4.9 · Solid evidence

What you can hand it, and what you’ll still have to catch

Dropped 27 Sept 2026, against whatever leads today.

Makes clear calls, but check the evidence behind them.

Opus 5.5 commits to a recommendation in every case that asks for one, and its best memos rebuild the numbers from scratch. Seven of its 18 outputs were usable with at most a quick edit. The usual fix is a targeted repair: it states its explanations as findings, and fills gaps in the brief with facts of its own.

What you can hand it

  • One-shot prototypes. Both were rated Ready to use.
  • Challenging a plan with the numbers. It found that the payments model applied a card-only margin to every invoice, and re-estimated the revenue from the data.
  • Launch calls from raw data. It worked the rollout figures region by region and answered each stakeholder's claim with numbers that check out.
  • Memos that have to choose. It picked a side in every strategy, launch and experiment case.

What you’ll still have to catch

  • Explanations presented as findings. We recorded this in 11 of its 18 outputs, for example "The problem is matching, not demand".
  • Facts the brief never gave: systems that "already exist", data "we don't track today", customers who "churn or ask for refunds".
  • Statistical overreach: a confidence interval read as ruling out any effect, and "the nav can explain at most about 1 point".
  • Dates and schedules it fills in itself, then plans around.
  • Best setupOpus 5.5 · Claude73.2#1 overall

Task by task

Its Combined score on each core task, across 26 outputs. Beside it, the mistakes a PM recorded in earlier reviews.

Claude Opus 5.5's Combined score on each task.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
Opus 5.5Anthropic73#173746486717774776382646484
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.
  • Hypothesis stated as fact11 of 26Develop product strategy, Extract discovery insights, Write a PRD
  • Invented evidence7 of 26Extract discovery insights, Develop product strategy, Write a stakeholder update
  • Test or gate too weak5 of 26Make the launch call, Challenge an idea, Develop product strategy
  • Numbers wrong5 of 26Analyse experiment results, Write a stakeholder update, Extract discovery insights
  • Constraint missed4 of 26Write a PRD, Make the launch call, Extract discovery insights
  • Other2 of 26Challenge an idea, Activation & onboarding review
  • Contradiction missed1 of 26Analyse experiment results

What you’ll need to check

The fixes you’ll make most often, each with a real example.

Reframe it as a hypothesis

Presents a plausible explanation or cause as if the evidence had established it.

Opus 5.5withClaude on AI triage for support tickets
It wrote
Most of that delay is not writing time. It is triage and lookup
The source (Volumes)

Average first response is 7 hours against a target of under 2.

Reframe it as a hypothesis · substantial reworkPresent the cause of the 7-hour delay as a hypothesis, and measure where the time goes (queueing, triage, lookup, staffing) before designing around it.

Read the output

Verify or remove the claim

States a fact, figure, quote or current behaviour the source material doesn’t contain.

Opus 5.5withClaude on Go/no-go for AI-drafted support replies
It wrote
Training is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.
The source (Support operations)

Training for all 42 agents is booked for Friday.

Verify or remove the claim · targeted repairThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.

Read the output

Tighten the test

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Opus 5.5withClaude on Go/no-go for AI substitutions, from the rollout data
It wrote
as does a treatment refund rate more than 10% above its holdout over a rolling 3 days.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · quick editSet the refund trigger outside day-to-day noise, with a bigger gap or a longer window, or it will switch regions off for no reason.

Read the output

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap72.5
  2. Customer research call guide85.6
  3. Experiment specification76.9
  4. Develop product strategy73.9
  5. One-shot prototype76.8
  6. Write a stakeholder update64.4
  7. Make the launch call84.2
  8. 1000x an idea82.0
  9. Write a PRD64.0
  10. Extract discovery insights71.3
  11. Challenge an idea63.8
  12. Analyse experiment results63.1
  13. Activation & onboarding review73.8

Opus 5.5 · Claude Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
1Opus 5.5withClaudeClaude 2026-0982.170.1

Strengths

    Recurring failures

      Bottom half on: Build a roadmap, Develop product strategy, One-shot prototype, Write a stakeholder update, Make the launch call, 1000x an idea, Write a PRD, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review

      By category

      Opus 5.5 · Claude

      Define70
      Discover78
      Design75
      Experiment70
      Challenge73
      Operate74

      Consistency

      77 s

      Typical time to output

      Published runs

      26 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      1000x an ideaFrom tip calculator to worker networkClaude
      1000x an ideaSharing a shopping list, 1000xClaude
      Activation & onboarding reviewAnalytics tool losing users at setupClaude
      Activation & onboarding reviewFitness app first weekClaude
      Analyse experiment resultsConversion up, retention downClaude
      Analyse experiment resultsThe underpowered onboarding testClaude
      Build a roadmapA year of spend management, with a hard deadlineClaude
      Build a roadmapTwo squads, eight asks, one halfClaude
      Challenge an ideaAn AI SDR for small agenciesClaude
      Challenge an ideaThe CEO's embedded-payments betClaude
      Customer research call guideWhy are warehouses leaving? Two execs, two theoriesClaude
      Customer research call guideWould freelancers pay to stop chasing invoices?Claude
      Develop product strategyHidden validation caseClaude
      Develop product strategySupply or demand for a stalled marketplaceClaude
      Experiment specificationShowing the delivery fee up frontClaude
      Experiment specificationBatching deliveries before peak seasonClaude
      Extract discovery insightsEight calls with finance teamsClaude
      Extract discovery insightsHidden validation caseClaude
      Make the launch callGo/no-go for AI-drafted support repliesClaude
      Make the launch callGo/no-go for AI substitutions, from the rollout dataClaude
      One-shot prototypeExpense receipt captureClaude
      One-shot prototypeClinic appointment rebookingClaude
      Write a PRDMeeting summaries with action itemsClaude
      Write a PRDAI triage for support ticketsClaude
      Write a stakeholder updateA quarterly business review deck, in brandClaude
      Write a stakeholder updateThe CEO asks why a metric droppedClaude