Ratings / Anthropic

Claude Sonnet 5.5

Anthropic · released 28 Sept 2026 · claude-sonnet-5-5 · effort high (default)

Sonnet 5.5withAPIProvisional
24 runs · ±5.3 · Solid evidence

What you can hand it, and what you’ll still have to catch

Dropped 30 Sept 2026, against Claude Sonnet 5 and whatever leads today.

Makes the right launch call. Double-check its facts.

Sonnet 5.5 is at its best when the job is a judgement call. It made the strongest launch calls of any model we've tested, and its challenges go straight for the assumption that matters. The catch is how it gets there: it often states an interpretation as fact, or uses a figure the brief never gave it. That's why the LLM judge found only 7 of its 13 graded outputs usable with a quick edit. Budget time to check its evidence.

What you can hand it

  • Launch calls. It weighs the evidence against the agreed criteria and lands on a clear recommendation you can defend.
  • Challenging an idea. It finds the load-bearing assumption (is growth really limited by leads?) and separates the real risks from the fixable ones.
  • Discovery synthesis. It treats what customers actually did as stronger evidence than what they said they'd do.
  • One-shot prototypes. The core flow worked, and it avoided the generic AI look.

What you’ll still have to catch

  • Interpretations stated as fact. A strategy memo asserted a £50 acquisition cost and "a search experience they find frustrating", with nothing in the brief behind either.
  • Numbers. It called a drop from 52% to 48% "4% of trial starters" when it's about 7.7%. Activation reviews were its weakest task.
  • Experiment readouts that skip the trust check. It never checked the sample ratio or logging before reading a result.
  • Length and ownership. It ran over the word limit twice, and its exec updates report misses without saying what the team has already done.
  • Best setupSonnet 5.5 · API80.9provisional
  • Against the leader+7.7Opus 5.5 · Claude

Task by task

Its Combined score on each core task, across 24 outputs. Beside it, the mistakes a PM recorded in earlier reviews.

Claude Sonnet 5.5's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
Sonnet 5.5Anthropic81Partial938476847893Best588780–826791Best
Opus 5.5Anthropic73#173746486717774776382646484
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

What you’ll need to check

The fixes you’ll make most often, each with a real example.

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap92.9leads
  2. Customer research call guide84.1
  3. Experiment specification86.5
  4. Develop product strategy83.5
  5. One-shot prototype92.9leads
  6. Write a stakeholder update67.0
  7. Make the launch call91.2leads
  8. 1000x an idea–
  9. Write a PRD75.6
  10. Extract discovery insights77.5
  11. Challenge an idea82.3
  12. Analyse experiment results79.7
  13. Activation & onboarding review57.5

Sonnet 5.5 · API Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
Prov.Sonnet 5.5withAPIOpenRouter 2026-0985.777.1

Strengths

    Leads: Build a roadmap, One-shot prototype, Make the launch call

    Recurring failures

      Bottom half on: Customer research call guide, Write a stakeholder update, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review

      By category

      Sonnet 5.5 · API

      Define84
      Discover81
      Design75
      Experiment83
      Challenge82
      Operate79

      Consistency

      57 s

      Typical API response time (app runs are timed by hand, so not comparable)

      Versus Claude Sonnet 5: not tested on this suite.

      Published runs

      24 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      Activation & onboarding reviewAnalytics tool losing users at setupAPI
      Activation & onboarding reviewFitness app first weekAPI
      Analyse experiment resultsConversion up, retention downAPI
      Analyse experiment resultsThe underpowered onboarding testAPI
      Build a roadmapTwo squads, eight asks, one halfAPI
      Build a roadmapA year of spend management, with a hard deadlineAPI
      Challenge an ideaAn AI SDR for small agenciesAPI
      Challenge an ideaThe CEO's embedded-payments betAPI
      Customer research call guideWould freelancers pay to stop chasing invoices?API
      Customer research call guideWhy are warehouses leaving? Two execs, two theoriesAPI
      Develop product strategySupply or demand for a stalled marketplaceAPI
      Develop product strategyHidden validation caseAPI
      Experiment specificationBatching deliveries before peak seasonAPI
      Experiment specificationShowing the delivery fee up frontAPI
      Extract discovery insightsEight calls with finance teamsAPI
      Extract discovery insightsHidden validation caseAPI
      Make the launch callGo/no-go for AI substitutions, from the rollout dataAPI
      Make the launch callGo/no-go for AI-drafted support repliesAPI
      One-shot prototypeExpense receipt captureAPI
      One-shot prototypeClinic appointment rebookingAPI
      Write a PRDAI triage for support ticketsAPI
      Write a PRDMeeting summaries with action itemsAPI
      Write a stakeholder updateThe CEO asks why a metric droppedAPI
      Write a stakeholder updateA quarterly business review deck, in brandAPI