Methodology

The full rules behind every number on this site. Most people want the short version.

Humanity’s Last PRD measures how these models and setups do on this suite of PM tasks. It isn’t a universal measure of product management ability.

Task-selection principles

Tasks are real PM jobs, the kind that turn up in a normal week. Each one has to pass three tests:

  • A decision is at stake. The output feeds a product decision or the way it gets communicated.
  • Good and bad are distinguishable. Two experienced PMs would mostly agree on which of two outputs is better.
  • Failure has a recognisable shape. Every task names the failure it's built to catch: invented consensus, causal overreach, scope creep.

We started with a small set of high-signal tasks. New ones join only when the current set shows its rubrics actually separate one setup from another.

Case sources

Cases are synthetic (written for the benchmark), anonymised from real work with identifying details changed, or built from public material. Each case records which. A small set is held back as a private validation set and never published, so results can be checked for training contamination.

Prompt and context policy

Every setup gets the same brief and the same supplied context, pasted verbatim. There's no per-model prompt tuning. When a harness adds its own instructions or skills, those are part of the setup and are listed on its record. Changing a brief, its context or its rubric creates a new case version, and old results stay attached to the version they ran against.

Model settings

Each model record stores the exact API identifier, the reasoning or effort setting, temperature where it's configurable, the context limit and a dated pricing snapshot. App harnesses (Claude, ChatGPT, Gemini) use the product defaults of the date shown.

Harness definitions

A harness is everything around the model: the environment (API, Claude, Claude Code, ChatGPT, Codex, Gemini), system instructions, skills such as No Surprises or Roast Me, tools and permissions. Results are always reported for a Model · Harness setup. Tasks that mostly measure the raw model are marked Measures the model. Tasks where tools and instructions do real work are marked Measures the system.

Number of repetitions

Most cases run once per setup. Cases where the result is noisy run three times. Where repeats exist, the spread across them is published as repeat spread.

Blind-review process

An experienced PM reads every output without knowing which model or harness produced it. The review queue hides identity until a score is submitted. The review answers one question:

Would we use this output to help make or communicate the product decision?

The reviewer also notes what it understood, what it missed and the decisive reason for the score. Only after submitting does the reviewer see the setup and the automated grades, and the original score can't be changed.

ScoreMeaning
1Misleading or unusable
2Substantial rework needed
3Useful starting point
4Strong, requires minor editing
5We'd use or send this

Automated grading

Structured checks come from a decision model: a grading model that answers each criterion as an atomic, typed question, evaluated independently against the brief, context and output. It returns a decision (pass, partial or fail) with a probability distribution and a confidence. The standard criteria are:

  • Uses the supplied evidence correctly
  • Addresses the actual decision
  • Respects explicit constraints
  • Identifies material uncertainty
  • Avoids unsupported claims
  • Produces the required deliverable
  • Doesn't commit the task's critical failures

Prototype tasks are also checked in a real browser. The prototype, delivered as a single HTML file, is opened in headless Chromium and its design is measured (colours, gradients, fonts, repeated card grids, em dashes). A small vision model from a provider we don't benchmark tries the task's goal and answers each browser check from what it sees. A stuck attempt counts as inconclusive, not as a failure, and those are decided by hand.

Written tasks are also graded by an independent LLM judge: a model from a family that isn't in the benchmark, which never sees which model or harness wrote the output. It first checks each claim the output makes about the current situation against the brief. Then it decides the same criteria against their pass, partial and fail definitions. Finally it answers the blind-review question on the same 1–5 scale. That answer, its reasoning and any claims it couldn't find in the brief are published beside each run. The PM reviewer only sees the judge's verdict after submitting their own. The judge was added after a trial in which it tracked the blind reviews more closely than the structured checks did.

Neither grader is treated as ground truth. Decisions below 65% confidence, and runs where the decision model and the PM review differ by more than 30 points, get a second look. We track agreement, false passes and false failures per criterion for both graders.

Scoring formula

Every run's component scores are published beside its PM score:

  • Decision model score. Pass counts 1, partial 0.5 and fail 0, averaged across criteria and scaled to 100.
  • LLM judge score. The judge's 1–5 answer to the blind-review question, mapped to 0–100 the same way as the PM review. Written tasks only.
  • PM review score. The blind reviewer's 1–5 judgement mapped to 0–100 (1 → 0, 3 → 50, 5 → 100).
  • PM score. 30% decision model, 35% LLM judge and 35% PM review. The LLM judge doesn't grade prototype tasks (they're checked in a browser instead), so there the decision model and the PM review split its share in proportion: about 46% and 54%.

Uncertainty is shown two ways. ± is the 95% margin on a setup's mean PM score across all its runs, left out below five runs. Repeat spread is the standard deviation between repeated runs of the same case. It measures consistency, not uncertainty about the mean.

Task scores average a setup's runs on that task. Overall and category scores average the task scores, so a task with more cases doesn't count for more.

PM score = 0.3 × decision model + 0.35 × LLM judge + 0.35 × PM review · capped at 40 on a critical failure

Treatment of failures and incomplete runs

Each case lists critical failures, such as recommending a launch that breaches a guardrail. A critical failure caps that run's PM score at 40, however good the rest of the writing is.

Only setups that have done every core task get an overall rank. Partial coverage shows as provisional, below the ranked list, and is never statistically adjusted to look comparable. Runs with failed or pending evaluations are never published.

Drops

There are no big versioned releases. When a new model comes out, it goes through the whole suite, every output is graded and read blind, and then its results go live together. That's a drop, dated and listed on the Drops page, with a verdict on whether it's worth switching to from the previous version and from whatever leads today.

After a drop, any new run of that model goes live as soon as it's been reviewed. Scores are always calculated live from the underlying grades. Rubric or brief changes create a new case version, and runs stay attached to the version they ran against.

Conflicts and limitations

There's one reviewer, an experienced PM, which makes this a single informed perspective, not a consensus. App harnesses are costed at API list prices, which overstates the marginal cost for subscribers. Model access is paid for out of pocket, and there's no sponsorship from model providers.