How we test

The short version. Everything below is in the full methodology, in more detail than most people need.

  • Real work, same briefEvery task is a job PMs actually do, with a realistic brief and messy inputs. Every model gets exactly the same brief and context. No per-model prompt tuning.
  • Setups, not just modelsOpus 5.5 in the Claude app and Opus 5.5 in Claude Code with a skill are different entries. When the harness does the work, the harness gets the credit.
  • Three gradersA decision model runs structured checks against the brief (30%). An independent LLM judge, from a model family we don’t benchmark, reads the written work and scores it (35%). An experienced PM reads every output blind and asks one question: would we use this to make or explain the decision? (35%).
  • Some mistakes sink a runRecommending a launch that breaks a guardrail, inventing evidence and the like cap the score at 40, however polished the rest reads.
  • Ranks are earnedA setup only gets an overall rank once it has done every core task. Anything less shows as provisional, and we don't adjust it to look comparable.
  • Drops, not hypeWhen a new model comes out, it goes through the whole suite before it shows up here. That's a drop: the results publish together, with a straight answer on whether it's worth switching.
  • What it isn'tOne suite of tasks and one reviewer. It's a well-informed view of how these models do PM work, not a measure of intelligence or of product management ability.

Full methodology