No overall pick yet
Nothing has finished the core suite yet
A model only gets an overall rank once it has a graded, blind-reviewed output on all 9 core tasks. Until then, partial results show below as provisional. Half a benchmark isn’t a verdict.
AI models tested on real PM work, so you can cut through the hype and decide what’s worth switching for.
0 models0 setups9 core tasks0 outputs read blind
A model only gets an overall rank once it has a graded, blind-reviewed output on all 9 core tasks. Until then, partial results show below as provisional. Half a benchmark isn’t a verdict.
Who leads each kind of PM work. Fewer outputs per category, so treat these as a lean, not a law.
Scores out of 100. Only setups that have done every core task get a rank.
Nothing published yet. Every output gets graded and read blind before it shows up here.