Which model is best at which PM task?

The only LLM benchmark that matters to product execs. AI models tested and evaluated on real product work across a broad set of tasks.

7 models13 core tasks182 outputs graded13 of 13 tasks calibrated

Combined score by model and task. Choose a column heading to sort by that task.
ModelDiscoverDesignDefineExperimentChallengeOperate
GPT-6.1 SolOpenAI89#194Best9389100Best87876988Best938898Best8088Best
GPT-6 AstraOpenAI88#28695Best7098888195Best7398Best96Best9588Best88
GPT-6 LunaOpenAI81#3618485907789Best67789282937979
Sonnet 5.5Anthropic80#4867893Best5895Best8476848074816387
Opus 5.5Anthropic72#574697771727464806281646484
Gemini 3.8 FlashGoogle52#673517656585238473748484147
Gemini 3.5 Flash-LiteGoogle45#738414959175843255337695044
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

Inside the benchmark

The short version of every page. Each one goes a lot deeper.

GPT-6.1 Sol leads by 1 point at a fifth of GPT-6 Astra’s price.

  1. GPT-6.1 Sol · API89
  2. GPT-6 Astra · ChatGPT88
  3. GPT-6 Luna · API81
  4. Sonnet 5.5 · API80
  5. Opus 5.5 · Claude72
Every setup, with filters

The check AI fails most: in customer research call guide, “marks what to cut if the call runs over” passes just 29% of the time.

  1. Marks what to cut if the call runs overCustomer research call guide29%
  2. Proposes tests that could failAcross 3 tasks36%
  3. Says how many sources support each findingExtract discovery insights46%
  4. Sized from the real trafficExperiment specification57%
What AI gets right and wrong, task by task

All 13 tasks are calibrated: the graders agree with our PM.

  • Customer research call guide
  • Extract discovery insights
  • One-shot prototype
  • Activation & onboarding review
  • Build a roadmap
  • Develop product strategy
  • Write a PRD
  • Experiment specification
  • Analyse experiment results
  • 1000x an idea
  • Challenge an idea
  • Write a stakeholder update
  • Make the launch call

Same briefTwo graders, one checklistA PM checks the checkers

How the scores work

Recent drops

All drops
Dropped

Gemini 3.8 Flash

Use it to find the angle, not to write the whole memo.

Gemini 3.8 Flash is good at spotting what a brief is really about. It found the mechanism hidden in the data on both 1000x cases, built roadmaps around their dependencies and a hard deadline, and wrote a customer call guide you could use with a quick edit. The trouble is everything around those insights: it fills gaps in the brief with facts of its own, and states its guesses as findings. That's why the LLM judge found only 1 of its 24 graded written outputs usable with a quick edit, and five of its 26 outputs made a critical mistake.

Gemini 3.8 Flash's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDiscoverDesignDefineExperimentChallengeOperate
Customer research call guideExtract discovery insightsOne-shot prototypeActivation & onboarding reviewBuild a roadmapDevelop product strategyWrite a PRDExperiment specificationAnalyse experiment results1000x an ideaChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6.1 SolOpenAI89#194Best9389100Best87876988Best938898Best8088Best
Gemini 3.8 FlashGoogle52#673517656585238473748484147
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.
What you can hand it, and what you’ll still have to catch
  • Best setupGemini 3.8 Flash · API51.7#6 overall
  • Against the leader−37.0GPT-6.1 Sol · API
Dropped

Claude Sonnet 5.5

Makes the right launch call. Double-check its facts.

Sonnet 5.5 is at its best when the job is a judgement call. It made the strongest launch calls of any model we've tested, and its challenges go straight for the assumption that matters. The catch is how it gets there: it often states an interpretation as fact, or uses a figure the brief never gave it. That's why the LLM judge found only 7 of its 13 graded outputs usable with a quick edit. Budget time to check its evidence.

Written 30 Sept 2026. We’ve added 4 tasks since (Build a roadmap, Customer research call guide, Experiment specification and 1000x an idea), so the scores here cover more than the write-up does.

Claude Sonnet 5.5's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDiscoverDesignDefineExperimentChallengeOperate
Customer research call guideExtract discovery insightsOne-shot prototypeActivation & onboarding reviewBuild a roadmapDevelop product strategyWrite a PRDExperiment specificationAnalyse experiment results1000x an ideaChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6.1 SolOpenAI89#194Best9389100Best87876988Best938898Best8088Best
Sonnet 5.5Anthropic80#4867893Best5895Best8476848074816387
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.
What you can hand it, and what you’ll still have to catch
  • Best setupSonnet 5.5 · API79.7#4 overall
  • Against the leader−8.9GPT-6.1 Sol · API
Dropped

GPT-6.1 Sol

Great at the thinking. Check the follow-through.

GPT-6.1 Sol is very good at the part of PM work that needs judgement. 15 of its 16 graded outputs are usable with at most a quick edit, and no model we've tested is better at challenging an idea. Where it slips is the follow-through: the rollback trigger on a launch call, the trade-off rule in a PRD, the "here's what we've already done" in an exec update.

Written 30 Sept 2026. We’ve added 4 tasks since (Build a roadmap, Customer research call guide, Experiment specification and 1000x an idea), so the scores here cover more than the write-up does.

GPT-6.1 Sol's Combined score on each task.
ModelOverallAll core tasksDiscoverDesignDefineExperimentChallengeOperate
Customer research call guideExtract discovery insightsOne-shot prototypeActivation & onboarding reviewBuild a roadmapDevelop product strategyWrite a PRDExperiment specificationAnalyse experiment results1000x an ideaChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6.1 SolOpenAI89#194Best9389100Best87876988Best938898Best8088Best
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.
What you can hand it, and what you’ll still have to catch
  • Best setupGPT-6.1 Sol · API88.7#1 overall