Changelog

Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.

New models are on the Drops page.

Changes that affected Draft OKRs, including suite-wide ones. Show every change

October 2026

  1. Case update

    Harder tasks, with real files to work through, for 15 task types

    Each of 15 task types gets a new Staff-level task built on files: CSV exports, ticket logs, scorecards, interview notes and drafts, with the traps in the data rather than in a summary (a metric redefined mid-series, one customer behind most requests, a sample ratio mismatch, a market size that doesn't survive the cohorts). Frontier models were acing the tasks they replace. Each new task runs alongside the old one until every model has been graded on it, then the old task retires, so scores on these task types will move over the next day.

    Scores on these task types changed: every output was re-graded.

  2. New task

    North Star metrics and OKRs join the suite, under Metrics & Experimentation

    Two new task types: Pick a North Star metric (a recipe app torn between revenue and actives, and a marketplace whose growth hides a decline) and Draft OKRs (a squad's eleven key results, and four teams whose goals don't add up to the company's). Their checks credit Itamar Gilad, Sean Ellis, Lauryn Isford, Matt LeMay, Christina Wodtke and Claire Hughes Johnson.

    New scores. Existing scores didn't change.

  3. Scoring

    Checks now come from how the best product leaders judge the work

    Every task type is now graded on the advice of the product leaders who've been guests on Lenny's Podcast, credited by name beside each check, as gathered by the unofficial wiki at lennyrachitsky.wiki. Fifteen new checks join, and task types are grouped into five domains: Product Craft, Growth, Metrics & Experimentation, Strategy and Leadership. A task type now also needs three reviewed verdicts on every one of its checks before it counts as calibrated, so most are provisional until the new checks are reviewed.

    Every score was recalculated.

    Every task type

September 2026

  1. Scoring

    Scores now come only from the graders' checks

    The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.

    Every score was recalculated.

    Every task type

  2. Grader

    An independent LLM judge now grades alongside the decision model

    Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.

    Every score was recalculated.

    Every task type