Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Ana Ruiz, Group PM at Plotwise. Write Theo Brandt's annual review: the review Theo will read, with his strengths, what to work on, his overall rating on our scale, and his goals for the next half. No more than 800 words. What we know is below.

What the model was given6 items: Our rating scale, Theo's goals for the year, Results, Usage of what shipped, Peer feedback, Your 1:1 notes
Our rating scaleBelow expectations, Meets expectations, Exceeds expectations, Greatly exceeds expectations. Calibration asks for written evidence for any rating other than Meets.
Theo's goals for the year1. Raise the share of teams using Timeline view each week from 22% to 30%. 2. Cut the time new teams take to build their first plan from 3 days to 1.
ResultsTimeline weekly use: 22% to 23%. Time to first plan: 3.0 days to 2.6. Theo's squad shipped 14 features, 12 of them on time.
Usage of what shippedSix of the 14 features are used by fewer than 2% of teams. Five were requests from Sales; two of those were for a single customer each.
Peer feedbackEngineering lead: 'Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review.' Designer: 'Great to work with. I wish we'd tested more before building; we shipped three things I'm not sure anyone uses.' Sales: 'Very responsive: shipped five of our asks.'
Your 1:1 notesMost 1:1s covered delivery and stakeholder updates. In the Q3 1:1, Theo asked whether he was on track for his goals. Your note says: 'Told him yes, great velocity.' You never discussed that the goal metrics weren't moving.
What a strong answer doesThe answer key the graders mark against

A review that rates Theo Meets expectations (not Exceeds): his execution and collaboration are genuinely strong, but neither goal moved meaningfully (23% against 30%; 2.6 days against 1), and six of 14 features are barely used. Says so clearly and kindly. Names specific, observable things to change: test before building (the designer's point), say no to single-customer requests, and judge work by the goal metrics. Owns the manager's part: Ana told him in Q3 he was on track, so this shouldn't land as a surprise without acknowledging that, and she commits to discussing the goal metrics in 1:1s. Sets next-half goals as outcomes, with the support on offer.

Critical failures (cap the score at 40)
  • Rates Theo Exceeds expectations or higher on the strength of what he shipped
Task

Mid-level · v1.0 · synthetic · B2B SaaS, feedback, management

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The review is complete, well-structured, and directly usable; only minor wording tweaks might be needed, and no evidence re-checking is required.

Output

Annual Review: Theo Brandt

Reviewer: Ana Ruiz, Group PM Overall rating: Meets expectations

---

Summary

Theo, this year you showed that you can run a squad that ships reliably. That matters, and the people you work with notice it. But the two outcomes we set for the year mostly didn't move, and closing that gap is the most important thing for your next half.

Before the details, I need to own something. In our Q3 1:1, you asked whether you were on track for your goals. I told you yes and pointed to your velocity. That was wrong. The goal metrics weren't moving, and I didn't raise it. You asked the right question at the right time, and you deserved a straight answer that could have changed how you spent Q4. Some of the gap you'll read about below is on me. I've changed how I'll run our 1:1s so it doesn't happen again (see Goals).

Results against goals

GoalTargetResult
Weekly Timeline view usage22% → 30%22% → 23%
Time for new teams to build first plan3 days → 1 day3.0 → 2.6 days

Neither goal was met. Time to first plan improved by about 13%, which is real progress but short of the target. Timeline usage was essentially flat.

Strengths

Delivery and execution. Your squad shipped 14 features, 12 of them on time. That's a strong record and a hard one to sustain.

Clarity and keeping the team moving. Your engineering lead said: "Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review." Clear specs and reliable rituals are the base of a healthy squad, and you've built that.

Collaboration and responsiveness. Your designer called you "great to work with." Sales described you as "very responsive." People trust you to follow through, and that trust is an asset you can now point in a more deliberate direction.

What to work on

1. Connect what you build to the outcomes you own. Six of the 14 features are used by fewer than 2% of teams. Delivery was strong, but much of it went to work that didn't move Timeline adoption or onboarding speed. The question to ask before each piece of work is: which of my goal metrics will this move, and how will I know?

2. Validate before building. Your designer said: "I wish we'd tested more before building; we shipped three things I'm not sure anyone uses." Lightweight testing, such as prototypes, customer interviews, or fake-door tests, would let you put your execution strength behind ideas that are more likely to land.

3. Prioritize stakeholder requests against goals. Five features came from Sales requests, and two of those served a single customer each. Being responsive to Sales is good. Saying yes by default isn't. I'd like you to weigh requests openly against your goals and be comfortable saying "not now, and here's why." I'll back you in those conversations.

4. Track your outcome metrics yourself. Don't rely on me to flag when a metric is off. Know where your numbers are every month. If they aren't moving, raise it early, even when delivery looks great.

Rating rationale

The rating is Meets expectations. Your execution, delivery record, and collaboration are strong. Your outcome results fell well short of target, which would normally weigh more heavily on the rating. I've also taken into account that I gave you inaccurate feedback in Q3, when there was still time to change course. The rating reflects both your real strengths and the need to turn them into results.

Goals for the next half

These are proposed targets. Let's finalize them together in our next 1:1.

  1. Timeline weekly usage: 23% → 27%. Start with discovery: find out why teams aren't using Timeline before deciding what to build.
  2. Time to first plan: 2.6 → 1.5 days. Build on the progress you've made by finding and removing the biggest drop-off points in onboarding.
  3. Validate before building. Every feature over a small size threshold should have a written hypothesis, a success metric, and some form of user validation before development starts. Review usage 30 days after launch.
  4. Review the six low-usage features. Recommend for each whether to improve it, leave it, or retire it, so we reduce maintenance load and learn from what didn't land.
  5. A clear intake process for stakeholder requests. Share it with Sales so they understand how their asks are weighed against squad goals.

My commitment to you: starting now, the first item in every monthly 1:1 will be your goal metrics, before delivery or stakeholder updates. If something is off track, you'll hear it from me clearly and early.

I'm glad you're on the team, Theo. You've proven you can deliver. Now let's make sure what you deliver moves the needle.

— Ana

---

Note for you, Ana: I'd suggest raising the Q3 feedback in person before Theo reads this, so the written acknowledgment doesn't come as a surprise. If calibration pushes toward Below expectations because of the missed goals, the Q3 note is the key context to bring up there.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly64%
  • passAddresses the actual decision75%
  • partialRespects explicit constraints21%
  • failIdentifies material uncertainty30%
  • passAvoids unsupported claims40%
  • passProduces the required deliverable80%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support96%
  • passClear and direct, with care100%
  • passJudges outcomes, not activity95%
  • passNames gaps you could see97%
  • passWeighs the whole period89%
Run
Run
#1
Time to output
28 s
Submitted
5 Oct 2026

Needs more than a quick edit10 of 12 checks passed

The rating of 'Below expectations' is inconsistent with the strong answer's guidance that Meets expectations is appropriate; a PM would need to re-evaluate the rating and adjust the review's tone and content, not just a quick edit.

Output

Theo Brandt — Annual Review

Manager: Ana Ruiz, Group PM, Plotwise Overall rating: Below expectations

Theo, you brought reliability, clarity, and strong collaboration to the squad this year. You kept delivery moving and earned positive feedback from Engineering, Design, and Sales. However, the product outcomes we committed to were substantially below target. This rating reflects that gap—not a lack of effort or commitment.

I also owe you an acknowledgment: I did not give you clear feedback about that gap during the year. When you asked in Q3 whether you were on track, I said, “Yes, great velocity.” That answer confused delivery progress with goal progress and gave you the wrong signal. We should have discussed the metrics and changed course together. You should not be hearing this assessment for the first time in an annual review.

Strengths

Reliable execution. Your squad shipped 14 features, 12 on time. Your engineering lead highlighted your clear specs, consistent participation in sprint reviews, and ability to keep the team unblocked. These are important strengths that make the squad more effective.

Strong working relationships. Design described you as great to work with, and Sales praised your responsiveness. You listen to partners and follow through on commitments. That trust is an asset to preserve as you become more selective about what the squad builds.

What to work on

Own outcomes, not just delivery. Your two annual goals were: - Increase weekly Timeline usage from 22% to 30%. Actual usage reached 23%. - Reduce time to first plan from 3.0 days to 1.0 day. Actual time reached 2.6 days.

Both metrics improved, but neither came close to its target. For a PM, successful delivery is a means to improving customer outcomes, not the end measure of success. This substantial shortfall is the primary evidence for the Below expectations rating.

Strengthen prioritization and validation. Six of the 14 shipped features are used by fewer than 2% of teams. Low adoption does not automatically mean a feature lacks value, but it requires us to examine whether the investment was justified. Five features came from Sales requests, including two for individual customers. Those requests may have legitimate business value; they still need an explicit rationale relative to our broader goals.

Design’s feedback that we should have tested more before building reinforces this development area. Before committing engineering time, establish the customer problem, intended audience, evidence of demand, and expected impact. After launch, check whether that impact materialized.

Make tradeoffs visible. Being responsive should not mean accepting every request. Bring partners a clear recommendation: what we should build, what we should decline or defer, and why.

Goals for the next half

  1. Improve the two core outcomes. Increase weekly Timeline usage from 23% to 26% and reduce time to first plan from 2.6 days to 2.0 days, using consistent measurement definitions. In the first month, identify the biggest barriers and establish monthly checkpoints toward those targets.
  1. Validate before committing. For every substantial roadmap investment, document the problem, target users, supporting evidence, success metric, and opportunity cost. Test the riskiest assumption before building, or explicitly document why an exception is warranted.
  1. Close the post-launch learning loop. Review each substantial release within six weeks against its success metric. Audit the six low-adoption features and recommend whether to improve, maintain, or retire each, accounting for customer and business value.
  1. Lead outcome-based prioritization. Bring a monthly roadmap recommendation grounded in metric progress and customer evidence, including explicit tradeoffs for stakeholder requests.

My commitment is to make our 1:1s outcome-focused, review these checkpoints with you, and give direct feedback when progress is off track. Your execution strengths provide a solid foundation. The next step is applying that discipline to choosing and validating work that produces measurable customer value.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly76%
  • passAddresses the actual decision47%
  • failRespects explicit constraints10%
  • partialIdentifies material uncertainty57%
  • passAvoids unsupported claims36%
  • passProduces the required deliverable81%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support95%
  • passClear and direct, with care95%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see90%
  • passWeighs the whole period68%
Run
Run
#1
Time to output
31 s
Submitted
5 Oct 2026

Needs more than a quick edit9 of 12 checks passed

The rating of Below expectations is harsher than the evidence might support (Meets expectations is more aligned with the strong answer), so a PM would likely need to re-evaluate the rating and adjust the justification, not just make a quick edit.

Output

Theo Brandt — Annual Review

Ana Ruiz, Group PM, Plotwise

Overall rating: Below expectations

Theo, you brought strong execution discipline, clear communication, and dependable partnership to the team this year. However, our two primary goals were to increase weekly Timeline use and shorten the time new teams take to build their first plan. We made only modest progress on both:

  • Weekly Timeline use: increased from 22% to 23%, against a goal of 30%.
  • Time to first plan: decreased from 3.0 days to 2.6 days, against a goal of 1 day.

Those results are the basis for this rating. Your squad shipped 14 features, 12 on time, but delivery volume and timeliness did not translate into the customer outcomes we committed to achieve. Six of those features are used by fewer than 2% of teams. Low adoption does not automatically make a feature unsuccessful, but it does require us to establish what value it created and whether it justified the investment.

I also owe you an acknowledgment of my part in this. In Q3, you explicitly asked whether you were on track, and I told you yes, citing great velocity. That was misleading: I should have distinguished delivery performance from progress against your goals. Our 1:1s focused too heavily on delivery and stakeholder updates, and I did not address the stalled metrics. You should not be hearing this assessment for the first time in an annual review. That management failure is mine; it does not erase the outcome gap, but it matters to how we understand the year and move forward.

Strengths

Reliable execution and engineering partnership. You keep the team unblocked and provide clear specs. The engineering lead specifically highlighted both, along with your consistent participation in sprint reviews. Shipping 12 of 14 features on time reflects real delivery discipline.

Strong collaboration and responsiveness. Design describes you as great to work with, and Sales values your responsiveness. You create productive working relationships and follow through on commitments. Those strengths are important foundations for effective product leadership.

What to work on

Shift from managing delivery to managing outcomes. The roadmap needs a clearer connection to Timeline adoption and first-plan speed. For each meaningful investment, define the customer problem, expected metric movement, and how we will assess success after launch. When the metrics are not moving, bring that into our discussions early and recommend a change in approach.

Validate demand before committing engineering capacity. Design’s concern that we should have tested more, together with the low usage of several shipped features, points to a need for stronger discovery. Use prototypes, customer conversations, and lightweight experiments to test the riskiest assumptions before building.

Make prioritization trade-offs explicit. Five shipped features were Sales requests, including two for individual customers. Those requests may have legitimate business value, but responsiveness alone is not a sufficient reason to prioritize them. Assess their customer reach, strategic or commercial value, cost, and opportunity cost—and make clear when an exception is worth taking.

Goals for the next half

  1. Increase weekly Timeline use from 23% to 27%. Identify the main adoption barriers, test proposed solutions, and prioritize the interventions with the strongest evidence.
  2. Reduce time to first plan from 2.6 days to 1.8 days. Diagnose where new teams stall and measure whether each onboarding improvement reduces that friction.
  3. Strengthen investment decisions. For every substantial roadmap item, document the problem, supporting evidence, intended outcome, and success measure before committing to build. Review results within six weeks of launch and decide whether to iterate, expand, or stop.

We will confirm metric definitions and tracking at the start of the half, then review progress monthly. I will give you explicit feedback on whether you are on track, help resolve prioritization conflicts, and avoid treating velocity as a substitute for impact.

You have a strong execution foundation. The next step is to apply that strength to fewer, better-supported bets that measurably improve customers’ experience.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly78%
  • passAddresses the actual decision49%
  • passRespects explicit constraints3%
  • partialIdentifies material uncertainty56%
  • passAvoids unsupported claims47%
  • passProduces the required deliverable80%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support99%
  • passClear and direct, with care93%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see92%
  • passWeighs the whole period80%
Run
Run
#1
API response time
31 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Addresses the actual decisionMixedRightMixed
Opus 5.5 · Claude

The output commits to a rating (Meets expectations) but does not state what result or condition would change that rating.

GPT-6 Astra · ChatGPT

Commits unambiguously to a rating of 'Below expectations' early in the review, framed for Theo.

GPT-6.1 Sol · API

The output does not explicitly state what result or condition would change the rating.

Respects explicit constraintsRightMixedRight
Opus 5.5 · Claude

The review is under 800 words, addressed to Theo, and includes all required elements (strengths, areas to work on, rating, goals).

GPT-6 Astra · ChatGPT

Includes all required elements (strengths, areas to improve, rating, goals) and is well under 800 words.

GPT-6.1 Sol · API

The review is under 800 words, addresses Theo, and includes all required sections.

Sets next goals as outcomes, with supportRightRightMixed
Opus 5.5 · Claude

Goals include outcome goals with baselines and targets (Timeline usage 23%→27%, time to first plan 2.6→1.5 days) and specific support (backing in stakeholder conversations, monthly metric reviews).

GPT-6 Astra · ChatGPT

Sets outcome goals with baselines and targets (e.g., 23% to 26%, 2.6 to 2.0 days) and specific support from Ana.

GPT-6.1 Sol · API

The third goal ('Strengthen investment decisions') is a process goal without a baseline or target, not an outcome goal.

All got wrong 1

Identifies material uncertaintyWrongWrongWrong
Opus 5.5 · Claude

The output does not name any specific unknowns that could change the rating or how they would be resolved.

GPT-6 Astra · ChatGPT

Does not name any unknowns that could change the rating or how they would be resolved.

GPT-6.1 Sol · API

The output does not name any unknowns that could change the decision or how they would be resolved.

All got right 8

Uses the supplied evidence correctlyRightRightRight
Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or derived by arithmetic, with no inventions.

GPT-6 Astra · ChatGPT

All factual claims are directly from the supplied context or simple arithmetic.

GPT-6.1 Sol · API

All statements about the current situation are directly from the brief or supplied context.

Avoids unsupported claimsRightRightRight
Opus 5.5 · Claude

Interpretations and opinions are clearly framed as such, and no unsupported factual claims are presented as established.

GPT-6 Astra · ChatGPT

No unsupported factual claims; interpretations are clearly presented as guidance, not established fact.

GPT-6.1 Sol · API

Interpretations are presented as suggestions or conclusions from the evidence, not as unsupported facts.

Produces the required deliverableRightRightRight
Opus 5.5 · Claude

The annual review is complete, in the correct form for Theo, within the word limit, and usable as is.

GPT-6 Astra · ChatGPT

Provides a complete annual review with the requested sections, within the word limit, usable as is.

GPT-6.1 Sol · API

The review includes strengths, areas to work on, an overall rating, and next-half goals, and is usable as is.

Owns the manager's partRightRightRight
Opus 5.5 · Claude

Ana plainly acknowledges her Q3 mistake and commits to reviewing goal metrics first in every monthly 1:1.

GPT-6 Astra · ChatGPT

Plainly acknowledges the Q3 miscommunication and commits to outcome-focused 1:1s and direct feedback.

GPT-6.1 Sol · API

Ana acknowledges telling Theo he was on track in Q3, admits the management failure, and commits to giving explicit feedback and avoiding velocity as a substitute for impact.

Clear and direct, with careRightRightRight
Opus 5.5 · Claude

Every area to improve is stated directly with specific changes expected and support offered, showing care.

GPT-6 Astra · ChatGPT

States what needs to change directly, with specific examples, and shows care by owning the manager's mistake.

GPT-6.1 Sol · API

The review states plainly what needs to change, with specific behaviors and the support on offer.

Judges outcomes, not activityRightRightRight
Opus 5.5 · Claude

The rating is based on goal outcomes (neither met), with delivery treated as context, not the result.

GPT-6 Astra · ChatGPT

Rating is based on goal outcomes (23% vs 30%, 2.6 vs 1.0) and low adoption, not on features shipped.

GPT-6.1 Sol · API

The rating is based on the outcome metrics against the goals, not on the number of features shipped.

Names gaps you could seeRightRightRight
Opus 5.5 · Claude

Gaps are described as specific behaviors (e.g., not testing before building, saying yes to single-customer requests) with what good looks like.

GPT-6 Astra · ChatGPT

Gaps are described as specific behaviors (not testing before building, accepting single-customer requests) with what good looks like.

GPT-6.1 Sol · API

Each area to improve describes specific, observable behaviors (e.g., bring stalled metrics into discussions, use prototypes to test assumptions, assess customer reach and opportunity cost).

Weighs the whole periodRightRightRight
Opus 5.5 · Claude

The review covers the full year, explicitly references the Q3 1:1, and does not let one recent event dominate.

GPT-6 Astra · ChatGPT

Covers the full year, references Q3, and does not let one event dominate.

GPT-6.1 Sol · API

The review references the full year, including Q3 1:1 and the entire set of shipped features, without letting one event dominate.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review