Changelog
Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.
New models are on the Drops page.
October 2026
Find the growth loop joins the suite
Can a model find a product's real growth loop, show whether it compounds, and pick the lever worth pulling? Two tasks: a form builder whose badge is the loop, and a Staff-level marketplace that's growing on the surface while its main loop decays. It's built on how Lenny's guests talk about growth loops.
New scores. Existing scores didn't change.
A clearer name for the call guide's time check
“Fits the call” is now “Marks what to cut if the call runs over”, because that's what outputs actually miss: their timings add up, but they don't say which questions to drop when time runs short. Only the name changed.
No change to scores.
September 2026
Scores now come only from the graders' checks
The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.
Every score was recalculated.
Every task type
Four new task types: roadmaps, customer call guides, experiment specs and 1000x
Each comes with two tasks, one mid-level and one Staff-level, with traps a strong answer has to find: a hard deadline, a sales promise the evidence doesn't back, contamination in a marketplace test, a CEO's idea that's just bigger adjectives. Every model has run all eight.
New scores. Existing scores didn't change.
A stricter “makes the call” check, and checks for the mistakes that kept recurring
“Addresses the actual decision” now only passes when an output commits to one answer and says what would change it. Five task types also gained a check drawn from the fixes reviewers kept making, such as getting the base of every number right, or saying how many sources support each finding.
Scores on these task types changed: every output was re-graded.
Every task type gets a practitioner check
Each task type now has one check for the sharpest thing experienced product leaders say about doing that job well, such as diagnosis before prescription for strategy, or trusting the data before reading an experiment. Every output was re-graded with it.
Scores on these task types changed: every output was re-graded.
Harder, Staff-level tasks for challenges, exec updates and launch calls
Each of the three task types swapped a task for a much harder one: a long, interlocking evidence pack, a slide deck in fixed brand colours, and a data workbook to work through. Scores on these task types now rest partly on the new tasks.
Scores on these task types changed: every output was re-graded.
An independent LLM judge now grades alongside the decision model
Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.
Every score was recalculated.
Every task type
Exec updates and launch calls join the suite
Two Operate task types join: writing a stakeholder update and making a launch call. “Sticking to the plan” leaves, because judging it fairly needs a real codebase, and every task type already penalises work nobody asked for.
New scores. Existing scores didn't change.
Prototypes are tested in a real browser, design included
One-shot prototypes are now opened in a headless browser, which tries the task's core flow and answers the browser checks. It also checks two signs of templated design: the stock AI palette, and grids of near-identical cards.
Scores on these task types changed: every output was re-graded.
The benchmark starts with eight task types
PRDs, strategy, discovery synthesis, one-shot prototypes, activation reviews, experiment readouts, challenging an idea and sticking to a plan, each graded check by check by the decision model.
New scores. Existing scores didn't change.