Changelog
Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.
New models are on the Drops page.
Changes that affected Find the growth loop, including suite-wide ones. Show every change
October 2026
Find the growth loop joins the suite
Can a model find a product's real growth loop, show whether it compounds, and pick the lever worth pulling? Two tasks: a form builder whose badge is the loop, and a Staff-level marketplace that's growing on the surface while its main loop decays. It's built on how Lenny's guests talk about growth loops.
New scores. Existing scores didn't change.
September 2026
Scores now come only from the graders' checks
The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.
Every score was recalculated.
Every task type
An independent LLM judge now grades alongside the decision model
Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.
Every score was recalculated.
Every task type