Changelog

Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.

New models are on the Drops page.

Changes that affected Find the growth loop, including suite-wide ones. Show every change

October 2026

  1. New task

    Find the growth loop joins the suite

    Can a model find a product's real growth loop, show whether it compounds, and pick the lever worth pulling? Two tasks: a form builder whose badge is the loop, and a Staff-level marketplace that's growing on the surface while its main loop decays. It's built on how Lenny's guests talk about growth loops.

    New scores. Existing scores didn't change.

September 2026

  1. Scoring

    Scores now come only from the graders' checks

    The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.

    Every score was recalculated.

    Every task type

  2. Grader

    An independent LLM judge now grades alongside the decision model

    Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.

    Every score was recalculated.

    Every task type