Changelog

Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.

New models are on the Drops page.

October 2026

  1. New task

    Find the growth loop joins the suite

    Can a model find a product's real growth loop, show whether it compounds, and pick the lever worth pulling? Two tasks: a form builder whose badge is the loop, and a Staff-level marketplace that's growing on the surface while its main loop decays. It's built on how Lenny's guests talk about growth loops.

    New scores. Existing scores didn't change.

  2. Checklist

    A clearer name for the call guide's time check

    “Fits the call” is now “Marks what to cut if the call runs over”, because that's what outputs actually miss: their timings add up, but they don't say which questions to drop when time runs short. Only the name changed.

    No change to scores.

September 2026

  1. Scoring

    Scores now come only from the graders' checks

    The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.

    Every score was recalculated.

    Every task type

  2. New task

    Four new task types: roadmaps, customer call guides, experiment specs and 1000x

    Each comes with two tasks, one mid-level and one Staff-level, with traps a strong answer has to find: a hard deadline, a sales promise the evidence doesn't back, contamination in a marketplace test, a CEO's idea that's just bigger adjectives. Every model has run all eight.

    New scores. Existing scores didn't change.

  3. Checklist

    A stricter “makes the call” check, and checks for the mistakes that kept recurring

    “Addresses the actual decision” now only passes when an output commits to one answer and says what would change it. Five task types also gained a check drawn from the fixes reviewers kept making, such as getting the base of every number right, or saying how many sources support each finding.

    Scores on these task types changed: every output was re-graded.

  4. Checklist

    Every task type gets a practitioner check

    Each task type now has one check for the sharpest thing experienced product leaders say about doing that job well, such as diagnosis before prescription for strategy, or trusting the data before reading an experiment. Every output was re-graded with it.

    Scores on these task types changed: every output was re-graded.

  5. Case update

    Harder, Staff-level tasks for challenges, exec updates and launch calls

    Each of the three task types swapped a task for a much harder one: a long, interlocking evidence pack, a slide deck in fixed brand colours, and a data workbook to work through. Scores on these task types now rest partly on the new tasks.

    Scores on these task types changed: every output was re-graded.

  6. Grader

    An independent LLM judge now grades alongside the decision model

    Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.

    Every score was recalculated.

    Every task type

  7. New task

    Exec updates and launch calls join the suite

    Two Operate task types join: writing a stakeholder update and making a launch call. “Sticking to the plan” leaves, because judging it fairly needs a real codebase, and every task type already penalises work nobody asked for.

    New scores. Existing scores didn't change.

  8. Grader

    Prototypes are tested in a real browser, design included

    One-shot prototypes are now opened in a headless browser, which tries the task's core flow and answers the browser checks. It also checks two signs of templated design: the stock AI palette, and grids of near-identical cards.

    Scores on these task types changed: every output was re-graded.

  9. Scoring

    The benchmark starts with eight task types

    PRDs, strategy, discovery synthesis, one-shot prototypes, activation reviews, experiment readouts, challenging an idea and sticking to a plan, each graded check by check by the decision model.

    New scores. Existing scores didn't change.