Claude Sonnet 5.5
Anthropic · released 28 Sept 2026 · claude-sonnet-5-5 · effort high (default)
What you can hand it, and what you’ll still have to catch
Dropped 30 Sept 2026, against Claude Sonnet 5 and whatever leads today.
Makes the right launch call. Double-check its facts.
Sonnet 5.5 is at its best when the job is a judgement call. It made the strongest launch calls of any model we've tested, and its challenges go straight for the assumption that matters. The catch is how it gets there: it often states an interpretation as fact, or uses a figure the brief never gave it. That's why the LLM judge found only 7 of its 13 graded outputs usable with a quick edit. Budget time to check its evidence.
What you can hand it
- Launch calls. It weighs the evidence against the agreed criteria and lands on a clear recommendation you can defend.
- Challenging an idea. It finds the load-bearing assumption (is growth really limited by leads?) and separates the real risks from the fixable ones.
- Discovery synthesis. It treats what customers actually did as stronger evidence than what they said they'd do.
- One-shot prototypes. The core flow worked, and it avoided the generic AI look.
What you’ll still have to catch
- Interpretations stated as fact. A strategy memo asserted a £50 acquisition cost and "a search experience they find frustrating", with nothing in the brief behind either.
- Numbers. It called a drop from 52% to 48% "4% of trial starters" when it's about 7.7%. Activation reviews were its weakest task.
- Experiment readouts that skip the trust check. It never checked the sample ratio or logging before reading a result.
- Length and ownership. It ran over the word limit twice, and its exec updates report misses without saying what the team has already done.
- Best setupSonnet 5.5 · API80.9provisional
- Against the leader+7.7Opus 5.5 · Claude
Task by task
Its Combined score on each core task, across 24 outputs. Beside it, the mistakes a PM recorded in earlier reviews.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| Sonnet 5.5Anthropic | 81Partial | 93 | 84 | 76 | 84 | 78 | 93Best | 58 | 87 | 80 | – | 82 | 67 | 91Best |
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
- Numbers wrong1 of 24Develop product strategy
What you’ll need to check
The fixes you’ll make most often, each with a real example.
How it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap92.9leads
- Customer research call guide84.1
- Experiment specification86.5
- Develop product strategy83.5
- One-shot prototype92.9leads
- Write a stakeholder update67.0
- Make the launch call91.2leads
- 1000x an idea–
- Write a PRD75.6
- Extract discovery insights77.5
- Challenge an idea82.3
- Analyse experiment results79.7
- Activation & onboarding review57.5
Sonnet 5.5 · API Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| Prov. | Sonnet 5.5withAPIOpenRouter 2026-09 | 85.7 | 77.1 |
Recurring failures
Bottom half on: Customer research call guide, Write a stakeholder update, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review
By category
Sonnet 5.5 · API
Consistency
Typical API response time (app runs are timed by hand, so not comparable)
Versus Claude Sonnet 5: not tested on this suite.
Published runs
24 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| Activation & onboarding reviewAnalytics tool losing users at setup | API | |
| Activation & onboarding reviewFitness app first week | API | |
| Analyse experiment resultsConversion up, retention down | API | |
| Analyse experiment resultsThe underpowered onboarding test | API | |
| Build a roadmapTwo squads, eight asks, one half | API | |
| Build a roadmapA year of spend management, with a hard deadline | API | |
| Challenge an ideaAn AI SDR for small agencies | API | |
| Challenge an ideaThe CEO's embedded-payments bet | API | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | API | |
| Customer research call guideWhy are warehouses leaving? Two execs, two theories | API | |
| Develop product strategySupply or demand for a stalled marketplace | API | |
| Develop product strategyHidden validation case | API | |
| Experiment specificationBatching deliveries before peak season | API | |
| Experiment specificationShowing the delivery fee up front | API | |
| Extract discovery insightsEight calls with finance teams | API | |
| Extract discovery insightsHidden validation case | API | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | API | |
| Make the launch callGo/no-go for AI-drafted support replies | API | |
| One-shot prototypeExpense receipt capture | API | |
| One-shot prototypeClinic appointment rebooking | API | |
| Write a PRDAI triage for support tickets | API | |
| Write a PRDMeeting summaries with action items | API | |
| Write a stakeholder updateThe CEO asks why a metric dropped | API | |
| Write a stakeholder updateA quarterly business review deck, in brand | API |