Changelog
Every change that can move a score: new task types, harder tasks, sharper checks, and changes to how we grade. If a number on the site moved, the reason is here.
New models are on the Drops page.
Changes that affected Write a performance review, including suite-wide ones. Show every change
October 2026
Harder tasks, with real files to work through, for 15 task types
Each of 15 task types gets a new Staff-level task built on files: CSV exports, ticket logs, scorecards, interview notes and drafts, with the traps in the data rather than in a summary (a metric redefined mid-series, one customer behind most requests, a sample ratio mismatch, a market size that doesn't survive the cohorts). Frontier models were acing the tasks they replace. Each new task runs alongside the old one until every model has been graded on it, then the old task retires, so scores on these task types will move over the next day.
Scores on these task types changed: every output was re-graded.
- Draft OKRs
- Hire a PM
- Write a performance review
- Experiment specification
- Analyse experiment results
- Extract discovery insights
- Build a roadmap
- Activation & onboarding review
- Write a stakeholder update
- Challenge an idea
- Customer research call guide
- Develop product strategy
- 1000x an idea
- Make the launch call
- Write a PRD
Hiring and performance reviews join the suite, under Leadership
Two Leadership task types join: Hire a PM (redesign a loop that keeps hiring for presence, and make the call on two finalists after a debrief swayed by the first speaker) and Write a performance review (a PM who shipped a lot but moved nothing, and a PIP proposed after one bad launch). Their checks credit Petra Wille and Kim Scott, and draw on the wiki's Hiring PMs and Performance reviews articles.
New scores. Existing scores didn't change.
Checks now come from how the best product leaders judge the work
Every task type is now graded on the advice of the product leaders who've been guests on Lenny's Podcast, credited by name beside each check, as gathered by the unofficial wiki at lennyrachitsky.wiki. Fifteen new checks join, and task types are grouped into five domains: Product Craft, Growth, Metrics & Experimentation, Strategy and Leadership. A task type now also needs three reviewed verdicts on every one of its checks before it counts as calibrated, so most are provisional until the new checks are reviewed.
Every score was recalculated.
Every task type
September 2026
Scores now come only from the graders' checks
The Combined score is now the share of checks an output passes, averaged across the decision model and the LLM judge, and capped at 40 for a critical failure. Our PM no longer scores outputs; they check the checkers. A task type counts as calibrated once the graders agree with them on at least 80% of checks.
Every score was recalculated.
Every task type
An independent LLM judge now grades alongside the decision model
Every written output is also graded by an LLM judge from a model family we don't test, so it isn't marking its own homework. It checks each claim against the brief, then answers every check yes or no.
Every score was recalculated.
Every task type