Ratings / OpenAI

GPT-6.1 Sol

OpenAI · released 29 Sept 2026 · gpt-6.1-sol · effort medium (default)

GPT-6.1 SolwithAPIProvisional
24 runs · ±3.8 · Solid evidence

What you can hand it, and what you’ll still have to catch

Dropped 30 Sept 2026, against GPT-6 Sol and whatever leads today.

Great at the thinking. Check the follow-through.

Sol is very good at the part of PM work that needs judgement. 15 of its 16 graded outputs are usable with at most a quick edit, and no model we've tested is better at challenging an idea. Where it slips is the follow-through: the rollback trigger on a launch call, the trade-off rule in a PRD, the "here's what we've already done" in an exec update. Sol ran through the API, so these results are the model on its own, without an app's tools or instructions.

What you can hand it

  • Challenging a plan. It goes straight for the assumption the whole plan rests on, and both of its challenges were usable with a quick edit.
  • Strategy memos. It names the real problem first (matching, not supply, was the constraint) and makes one clear bet instead of hedging across three.
  • Activation reviews. It defines activation by what predicts retention, then ranks fixes by impact, not by where they sit in the funnel.
  • Experiment readouts and discovery synthesis. Every one was usable with a quick edit.

What you’ll still have to catch

  • The safety net on launch calls. It makes the right call, then doesn't say what would make you roll it back, or what can't be undone if it's wrong.
  • PRD trade-offs and tests. It sets targets but not which goal wins when precision and coverage conflict, and some launch gates have no measurement window.
  • Exec updates that don't own the misses. It reports what went wrong without saying what the team has already done about it, and once left a note to itself in the deck: "Use Finance's figure, not CS's 102%".
  • Data trust. It read an experiment result without first checking the sample split.
  • Best setupGPT-6.1 Sol · API88.7provisional
  • Against the leader+15.5Opus 5.5 · Claude

Task by task

Its Combined score on each core task, across 24 outputs.

GPT-6.1 Sol's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6.1 SolOpenAI89Partial898769939389100Best8890–98Best8188
Opus 5.5Anthropic73#173746486717774776382646484
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap89.4
  2. Customer research call guide92.7
  3. Experiment specification87.5leads
  4. Develop product strategy87.2
  5. One-shot prototype88.7
  6. Write a stakeholder update81.2
  7. Make the launch call88.3
  8. 1000x an idea–
  9. Write a PRD68.9
  10. Extract discovery insights92.5
  11. Challenge an idea97.9leads
  12. Analyse experiment results90.5
  13. Activation & onboarding review100.0leads

GPT-6.1 Sol · API Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
Prov.GPT-6.1 SolwithAPIOpenRouter 2026-0989.987.4

Strengths

    Leads: Experiment specification, Challenge an idea, Activation & onboarding review

    Recurring failures

      By category

      GPT-6.1 Sol · API

      Define82
      Discover93
      Design94
      Experiment89
      Challenge98
      Operate85

      Consistency

      46 s

      Typical API response time (app runs are timed by hand, so not comparable)

      Versus GPT-6 Sol: not tested on this suite.

      Published runs

      24 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      Activation & onboarding reviewAnalytics tool losing users at setupAPI
      Activation & onboarding reviewFitness app first weekAPI
      Analyse experiment resultsConversion up, retention downAPI
      Analyse experiment resultsThe underpowered onboarding testAPI
      Build a roadmapA year of spend management, with a hard deadlineAPI
      Build a roadmapTwo squads, eight asks, one halfAPI
      Challenge an ideaAn AI SDR for small agenciesAPI
      Challenge an ideaThe CEO's embedded-payments betAPI
      Customer research call guideWould freelancers pay to stop chasing invoices?API
      Customer research call guideWhy are warehouses leaving? Two execs, two theoriesAPI
      Develop product strategyHidden validation caseAPI
      Develop product strategySupply or demand for a stalled marketplaceAPI
      Experiment specificationBatching deliveries before peak seasonAPI
      Experiment specificationShowing the delivery fee up frontAPI
      Extract discovery insightsEight calls with finance teamsAPI
      Extract discovery insightsHidden validation caseAPI
      Make the launch callGo/no-go for AI substitutions, from the rollout dataAPI
      Make the launch callGo/no-go for AI-drafted support repliesAPI
      One-shot prototypeClinic appointment rebookingAPI
      One-shot prototypeExpense receipt captureAPI
      Write a PRDAI triage for support ticketsAPI
      Write a PRDMeeting summaries with action itemsAPI
      Write a stakeholder updateThe CEO asks why a metric droppedAPI
      Write a stakeholder updateA quarterly business review deck, in brandAPI