GPT-6.1 Sol
OpenAI · released 29 Sept 2026 · gpt-6.1-sol · effort medium (default)
What you can hand it, and what you’ll still have to catch
Dropped 30 Sept 2026, against GPT-6 Sol and whatever leads today.
Great at the thinking. Check the follow-through.
Sol is very good at the part of PM work that needs judgement. 15 of its 16 graded outputs are usable with at most a quick edit, and no model we've tested is better at challenging an idea. Where it slips is the follow-through: the rollback trigger on a launch call, the trade-off rule in a PRD, the "here's what we've already done" in an exec update. Sol ran through the API, so these results are the model on its own, without an app's tools or instructions.
What you can hand it
- Challenging a plan. It goes straight for the assumption the whole plan rests on, and both of its challenges were usable with a quick edit.
- Strategy memos. It names the real problem first (matching, not supply, was the constraint) and makes one clear bet instead of hedging across three.
- Activation reviews. It defines activation by what predicts retention, then ranks fixes by impact, not by where they sit in the funnel.
- Experiment readouts and discovery synthesis. Every one was usable with a quick edit.
What you’ll still have to catch
- The safety net on launch calls. It makes the right call, then doesn't say what would make you roll it back, or what can't be undone if it's wrong.
- PRD trade-offs and tests. It sets targets but not which goal wins when precision and coverage conflict, and some launch gates have no measurement window.
- Exec updates that don't own the misses. It reports what went wrong without saying what the team has already done about it, and once left a note to itself in the deck: "Use Finance's figure, not CS's 102%".
- Data trust. It read an experiment result without first checking the sample split.
- Best setupGPT-6.1 Sol · API88.7provisional
- Against the leader+15.5Opus 5.5 · Claude
Task by task
Its Combined score on each core task, across 24 outputs.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6.1 SolOpenAI | 89Partial | 89 | 87 | 69 | 93 | 93 | 89 | 100Best | 88 | 90 | – | 98Best | 81 | 88 |
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
How it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap89.4
- Customer research call guide92.7
- Experiment specification87.5leads
- Develop product strategy87.2
- One-shot prototype88.7
- Write a stakeholder update81.2
- Make the launch call88.3
- 1000x an idea–
- Write a PRD68.9
- Extract discovery insights92.5
- Challenge an idea97.9leads
- Analyse experiment results90.5
- Activation & onboarding review100.0leads
GPT-6.1 Sol · API Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| Prov. | GPT-6.1 SolwithAPIOpenRouter 2026-09 | 89.9 | 87.4 |
Recurring failures
By category
GPT-6.1 Sol · API
Consistency
Typical API response time (app runs are timed by hand, so not comparable)
Versus GPT-6 Sol: not tested on this suite.
Published runs
24 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| Activation & onboarding reviewAnalytics tool losing users at setup | API | |
| Activation & onboarding reviewFitness app first week | API | |
| Analyse experiment resultsConversion up, retention down | API | |
| Analyse experiment resultsThe underpowered onboarding test | API | |
| Build a roadmapA year of spend management, with a hard deadline | API | |
| Build a roadmapTwo squads, eight asks, one half | API | |
| Challenge an ideaAn AI SDR for small agencies | API | |
| Challenge an ideaThe CEO's embedded-payments bet | API | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | API | |
| Customer research call guideWhy are warehouses leaving? Two execs, two theories | API | |
| Develop product strategyHidden validation case | API | |
| Develop product strategySupply or demand for a stalled marketplace | API | |
| Experiment specificationBatching deliveries before peak season | API | |
| Experiment specificationShowing the delivery fee up front | API | |
| Extract discovery insightsEight calls with finance teams | API | |
| Extract discovery insightsHidden validation case | API | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | API | |
| Make the launch callGo/no-go for AI-drafted support replies | API | |
| One-shot prototypeClinic appointment rebooking | API | |
| One-shot prototypeExpense receipt capture | API | |
| Write a PRDAI triage for support tickets | API | |
| Write a PRDMeeting summaries with action items | API | |
| Write a stakeholder updateThe CEO asks why a metric dropped | API | |
| Write a stakeholder updateA quarterly business review deck, in brand | API |