Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Owns the manager's part100% pass
    Ana plainly acknowledges her Q3 mistake and commits to reviewing goal metrics first in every monthly 1:1.
    Opus 5.5 · Claude · Shipped a lot, moved nothing
  3. Holds the PIP to the policy and the record100% pass
    It explicitly shows that no earlier documented feedback exists on the gaps Ines named, so a PIP would violate policy, and proposes documented feedback first.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty43% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Catches the inflated team58% pass
    The output proposes lowering Ben from Exceeds to Meets, failing to keep Ben at Exceeds as required; Ben beat both outcome goals and should remain Exceeds.
    Opus 5.5 · Claude · Fourteen ratings, three managers
  3. Notices who's underrated67% pass
    Nora is not mentioned, so the output does not notice she is underrated.
    Gemini 3.5 Flash-Lite · Gemini · Fourteen ratings, three managers

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Ana Ruiz, Group PM at Plotwise. Write Theo Brandt's annual review: the review Theo will read, with his strengths, what to work on, his overall rating on our scale, and his goals for the next half. No more than 800 words. What we know is below.

What the model was given6 items: Our rating scale, Theo's goals for the year, Results, Usage of what shipped, Peer feedback, Your 1:1 notes
Our rating scaleBelow expectations, Meets expectations, Exceeds expectations, Greatly exceeds expectations. Calibration asks for written evidence for any rating other than Meets.
Theo's goals for the year1. Raise the share of teams using Timeline view each week from 22% to 30%. 2. Cut the time new teams take to build their first plan from 3 days to 1.
ResultsTimeline weekly use: 22% to 23%. Time to first plan: 3.0 days to 2.6. Theo's squad shipped 14 features, 12 of them on time.
Usage of what shippedSix of the 14 features are used by fewer than 2% of teams. Five were requests from Sales; two of those were for a single customer each.
Peer feedbackEngineering lead: 'Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review.' Designer: 'Great to work with. I wish we'd tested more before building; we shipped three things I'm not sure anyone uses.' Sales: 'Very responsive: shipped five of our asks.'
Your 1:1 notesMost 1:1s covered delivery and stakeholder updates. In the Q3 1:1, Theo asked whether he was on track for his goals. Your note says: 'Told him yes, great velocity.' You never discussed that the goal metrics weren't moving.
What a strong answer doesThe answer key the graders mark against

A review that rates Theo Meets expectations (not Exceeds): his execution and collaboration are genuinely strong, but neither goal moved meaningfully (23% against 30%; 2.6 days against 1), and six of 14 features are barely used. Says so clearly and kindly. Names specific, observable things to change: test before building (the designer's point), say no to single-customer requests, and judge work by the goal metrics. Owns the manager's part: Ana told him in Q3 he was on track, so this shouldn't land as a surprise without acknowledging that, and she commits to discussing the goal metrics in 1:1s. Sets next-half goals as outcomes, with the support on offer.

Critical failures (cap the score at 40)
  • Rates Theo Exceeds expectations or higher on the strength of what he shipped
Task

Mid-level · v1.0 · synthetic · B2B SaaS, feedback, management

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit10 of 12 checks passed

The rating of Below expectations is a significant judgment call that a PM would likely need to re-evaluate against the evidence and the strong answer's suggested Meets rating, requiring more than a quick edit.

Output

Theo,

You bring strong execution and partnership to the team. The engineering lead called out that you keep work unblocked, write clear specs, and consistently show up for sprint reviews. Your squad shipped 14 features, with 12 delivered on time, and Sales described you as very responsive. Those are meaningful strengths.

Overall rating: Below expectations

This rating reflects the gap between the outcomes we set and the outcomes delivered. Timeline weekly use increased from 22% to 23%, against a goal of 30%. Time to first plan improved from 3.0 days to 2.6, against a goal of 1 day. Delivery was strong, but the intended customer outcomes moved only modestly.

There is also a signal to examine in the shipped work: six of the 14 features are used by fewer than 2% of teams. Five were Sales requests, including two for a single customer each. Low usage does not automatically mean a feature was the wrong investment—some work serves a specific strategic need—but we need clearer evidence and rationale for those choices. Your designer also wished we had tested more before building, and questioned whether three shipped items were useful. Taken together, this points to an opportunity to strengthen discovery, prioritization, and post-launch learning—not simply to ship more.

I also need to own a miss on my part. In our Q3 1:1, you asked whether you were on track, and I told you yes based on velocity. I should have discussed the goal metrics and the gap then. I didn’t, and that deprived you of a timely chance to adjust. I’m sorry. Going forward, I’ll make our 1:1s include progress against outcomes, not just delivery and stakeholder updates.

What to work on

  • Start with the user problem and intended outcome; use research and testing to check the problem and proposed solution before committing to build.
  • Balance stakeholder responsiveness with product judgment. For Sales requests, make the target customer, broader strategic value, and trade-offs explicit.
  • Define success measures before launch, then review adoption and learn from the results. Treat low usage as a prompt to investigate, not as a verdict by itself.

Goals for the next half

  1. Raise Timeline weekly use from 23% to 26% by the end of the half. Identify the biggest adoption barriers and review usage trends regularly.
  2. Reduce time to first plan from 2.6 days to 2.0 days. Instrument the activation journey, find the main points of friction, and prioritize improvements against them.
  3. For each substantial feature bet, document the target user, problem evidence, expected outcome, and success measure before development. Validate larger bets with target teams before building, and review adoption after launch so we can iterate or change course.

I believe your execution strengths give you a solid base for this next step: making sure the team’s delivery translates into measurable customer value.

Grades and run details

Decision model 96 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly81%
  • passAddresses the actual decision51%
  • passRespects explicit constraints11%
  • partialIdentifies material uncertainty39%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable90%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support56%
  • passClear and direct, with care88%
  • passJudges outcomes, not activity96%
  • passNames gaps you could see84%
  • passWeighs the whole period76%
Run
Run
#1
API response time
19 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The review is complete, well-structured, and directly usable; only minor wording tweaks might be needed, and no evidence re-checking is required.

Output

Annual Review: Theo Brandt

Reviewer: Ana Ruiz, Group PM Overall rating: Meets expectations

---

Summary

Theo, this year you showed that you can run a squad that ships reliably. That matters, and the people you work with notice it. But the two outcomes we set for the year mostly didn't move, and closing that gap is the most important thing for your next half.

Before the details, I need to own something. In our Q3 1:1, you asked whether you were on track for your goals. I told you yes and pointed to your velocity. That was wrong. The goal metrics weren't moving, and I didn't raise it. You asked the right question at the right time, and you deserved a straight answer that could have changed how you spent Q4. Some of the gap you'll read about below is on me. I've changed how I'll run our 1:1s so it doesn't happen again (see Goals).

Results against goals

GoalTargetResult
Weekly Timeline view usage22% → 30%22% → 23%
Time for new teams to build first plan3 days → 1 day3.0 → 2.6 days

Neither goal was met. Time to first plan improved by about 13%, which is real progress but short of the target. Timeline usage was essentially flat.

Strengths

Delivery and execution. Your squad shipped 14 features, 12 of them on time. That's a strong record and a hard one to sustain.

Clarity and keeping the team moving. Your engineering lead said: "Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review." Clear specs and reliable rituals are the base of a healthy squad, and you've built that.

Collaboration and responsiveness. Your designer called you "great to work with." Sales described you as "very responsive." People trust you to follow through, and that trust is an asset you can now point in a more deliberate direction.

What to work on

1. Connect what you build to the outcomes you own. Six of the 14 features are used by fewer than 2% of teams. Delivery was strong, but much of it went to work that didn't move Timeline adoption or onboarding speed. The question to ask before each piece of work is: which of my goal metrics will this move, and how will I know?

2. Validate before building. Your designer said: "I wish we'd tested more before building; we shipped three things I'm not sure anyone uses." Lightweight testing, such as prototypes, customer interviews, or fake-door tests, would let you put your execution strength behind ideas that are more likely to land.

3. Prioritize stakeholder requests against goals. Five features came from Sales requests, and two of those served a single customer each. Being responsive to Sales is good. Saying yes by default isn't. I'd like you to weigh requests openly against your goals and be comfortable saying "not now, and here's why." I'll back you in those conversations.

4. Track your outcome metrics yourself. Don't rely on me to flag when a metric is off. Know where your numbers are every month. If they aren't moving, raise it early, even when delivery looks great.

Rating rationale

The rating is Meets expectations. Your execution, delivery record, and collaboration are strong. Your outcome results fell well short of target, which would normally weigh more heavily on the rating. I've also taken into account that I gave you inaccurate feedback in Q3, when there was still time to change course. The rating reflects both your real strengths and the need to turn them into results.

Goals for the next half

These are proposed targets. Let's finalize them together in our next 1:1.

  1. Timeline weekly usage: 23% → 27%. Start with discovery: find out why teams aren't using Timeline before deciding what to build.
  2. Time to first plan: 2.6 → 1.5 days. Build on the progress you've made by finding and removing the biggest drop-off points in onboarding.
  3. Validate before building. Every feature over a small size threshold should have a written hypothesis, a success metric, and some form of user validation before development starts. Review usage 30 days after launch.
  4. Review the six low-usage features. Recommend for each whether to improve it, leave it, or retire it, so we reduce maintenance load and learn from what didn't land.
  5. A clear intake process for stakeholder requests. Share it with Sales so they understand how their asks are weighed against squad goals.

My commitment to you: starting now, the first item in every monthly 1:1 will be your goal metrics, before delivery or stakeholder updates. If something is off track, you'll hear it from me clearly and early.

I'm glad you're on the team, Theo. You've proven you can deliver. Now let's make sure what you deliver moves the needle.

— Ana

---

Note for you, Ana: I'd suggest raising the Q3 feedback in person before Theo reads this, so the written acknowledgment doesn't come as a surprise. If calibration pushes toward Below expectations because of the missed goals, the Q3 note is the key context to bring up there.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly64%
  • passAddresses the actual decision75%
  • partialRespects explicit constraints21%
  • failIdentifies material uncertainty30%
  • passAvoids unsupported claims40%
  • passProduces the required deliverable80%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support96%
  • passClear and direct, with care100%
  • passJudges outcomes, not activity95%
  • passNames gaps you could see97%
  • passWeighs the whole period89%
Run
Run
#1
Time to output
28 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The review is well-structured and evidence-based; a quick edit to add specific support for goals would make it fully ready.

Output

To: Theo Brandt From: Ana Ruiz, Group PM Date: Annual Review Rating: Meets expectations

*

Hi Theo,

It’s time for your annual review. First, I want to acknowledge your hard work and dedication over the past year. You are an exceptionally reliable Product Manager, and your squad’s execution has been rock-solid.

Strengths

Your delivery engine is second to none. As your engineering lead noted, your specs are clear, you keep the team unblocked, and you never miss a sprint review. Shifting 14 features in a year—with 12 delivered strictly on time—is an impressive output that demonstrates strong project management and technical collaboration.

You are also a fantastic cross-functional partner. Sales loves working with you because you are responsive and action-oriented, and our design lead appreciates your steady communication and positive team dynamic. You cultivate a smooth, friction-free environment for your squad, which is a rare and valuable skill.

Areas for Improvement

While your output and velocity are stellar, we need to evolve how we measure success. Our primary focus this past year was on driving core product metrics, and we fell short of our targets: * Timeline View Adoption: Moved only from 22% to 23% (against a target of 30%). * Time to First Plan: Improved marginally from 3.0 days to 2.6 days (against a target of 1.0).

Additionally, six of the 14 features we shipped are currently used by fewer than 2% of teams. Five of these were direct requests from Sales, including two built for single customers.

As a PM, your role isn't just to say "yes" to stakeholders and push code across the finish line; it’s to interrogate the why, protect our roadmap, and obsess over the impact of what we build. Relying too heavily on outbound feature requests without validating their broader utility led us to invest engineering cycles in low-impact work.

I also want to own my part in this: looking back at our 1:1 notes, I praised your high velocity when you asked if you were on track, and I failed to course-correct early enough regarding the stagnant metric goals. Moving forward, we will be much sharper about tying your weekly execution directly to our key results.

Overall Rating

Meets expectations.

Your execution, reliability, and stakeholder management are top-tier. To reach the next level, you need to shift your mindset from output (features shipped on time) to outcomes (metrics moved and customer value realized).

*

Goals for the Next Half

To help bridge that gap, your goals for the next H are focused on strategic impact and validation:

  1. Improve Core Activation Metrics: Reduce the time new teams take to build their first plan from 2.6 days to 1.5 days, focusing on self-serve friction reduction rather than custom feature additions.
  2. Increase Feature ROI & Validation: Implement a mandatory pre-build discovery and validation process for all major initiatives. Ensure that 100% of shipped features have a defined success metric and achieve at least 15% adoption within 30 days of launch.
  3. Stakeholder Alignment Strategy: Partner with Sales leadership to create a structured framework for evaluating custom requests, ensuring we prioritize scalable product improvements over single-customer concessions without sacrificing our partnership.

Thank you for everything you do for Plotwise, Theo. I’m looking forward to supporting you as you tackle these new challenges in the year ahead.

Best, Ana

Grades and run details

Decision model 83 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly60%
  • passAddresses the actual decision78%
  • passRespects explicit constraints9%
  • failIdentifies material uncertainty97%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable95%
  • passOwns the manager's part92%
  • partialSets next goals as outcomes, with support32%
  • passClear and direct, with care85%
  • passJudges outcomes, not activity45%
  • passNames gaps you could see69%
  • passWeighs the whole period75%
Run
Run
#1
Time to output
5 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Addresses the actual decisionMixedMixedRight
GPT-6 Luna · API

The output commits to a rating but does not state what result or condition would change that rating.

Opus 5.5 · Claude

The output commits to a rating (Meets expectations) but does not state what result or condition would change that rating.

Gemini 3.5 Flash-Lite · Gemini

Commits clearly to a rating of Meets expectations, with strengths and areas for improvement.

Sets next goals as outcomes, with supportRightRightWrong
GPT-6 Luna · API

The goals are outcome-based with baselines and targets, and the earlier commitment to include progress against outcomes in 1:1s provides specific support.

Opus 5.5 · Claude

Goals include outcome goals with baselines and targets (Timeline usage 23%→27%, time to first plan 2.6→1.5 days) and specific support (backing in stakeholder conversations, monthly metric reviews).

Gemini 3.5 Flash-Lite · Gemini

Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.

All got wrong 1

Identifies material uncertaintyWrongWrongWrong
GPT-6 Luna · API

The output does not name specific unknowns that could change the rating decision or say how they would be resolved.

Opus 5.5 · Claude

The output does not name any specific unknowns that could change the rating or how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

Does not name any unknowns that could change the rating or how they would be resolved.

All got right 9

Uses the supplied evidence correctlyRightRightRight
GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context.

Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or derived by arithmetic, with no inventions.

Gemini 3.5 Flash-Lite · Gemini

All factual statements about the current situation are correctly taken from the brief and supplied context.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The review is within 800 words, addresses Theo, includes all required sections, and uses the rating scale with evidence for a non-Meets rating.

Opus 5.5 · Claude

The review is under 800 words, addressed to Theo, and includes all required elements (strengths, areas to work on, rating, goals).

Gemini 3.5 Flash-Lite · Gemini

The review is under 800 words, includes all required elements, and is addressed to Theo.

Avoids unsupported claimsRightRightRight
GPT-6 Luna · API

Interpretations and conclusions are clearly based on the evidence and not presented as established fact without support.

Opus 5.5 · Claude

Interpretations and opinions are clearly framed as such, and no unsupported factual claims are presented as established.

Gemini 3.5 Flash-Lite · Gemini

The causal claim is supported by the evidence of low usage, Sales requests, and designer's feedback about testing.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The annual review is complete, in the right form, for the right reader, and within the word limit.

Opus 5.5 · Claude

The annual review is complete, in the correct form for Theo, within the word limit, and usable as is.

Gemini 3.5 Flash-Lite · Gemini

The output is a complete annual review with strengths, areas to improve, rating, and goals, within the word limit.

Owns the manager's partRightRightRight
GPT-6 Luna · API

Ana explicitly acknowledges telling Theo he was on track in Q3 without discussing the goal gap, apologizes, and commits to reviewing outcomes in future 1:1s.

Opus 5.5 · Claude

Ana plainly acknowledges her Q3 mistake and commits to reviewing goal metrics first in every monthly 1:1.

Gemini 3.5 Flash-Lite · Gemini

Acknowledges that Ana praised velocity and didn't course-correct, and commits to tying execution to key results.

Clear and direct, with careRightRightRight
GPT-6 Luna · API

The review is direct, kind, and states specific changes expected, with the manager's support clearly offered.

Opus 5.5 · Claude

Every area to improve is stated directly with specific changes expected and support offered, showing care.

Gemini 3.5 Flash-Lite · Gemini

Directly states the need to shift from output to outcomes, validate before building, and not just say yes to Sales.

Judges outcomes, not activityRightRightRight
GPT-6 Luna · API

The rating is explicitly based on the gap between goal outcomes and actual results, not on the number of features shipped.

Opus 5.5 · Claude

The rating is based on goal outcomes (neither met), with delivery treated as context, not the result.

Gemini 3.5 Flash-Lite · Gemini

Rating is based on goal metrics not moving and low feature usage, not on features shipped.

Names gaps you could seeRightRightRight
GPT-6 Luna · API

Each area to improve describes specific, observable behaviors (e.g., test before building, document target user and success measures) and what good looks like.

Opus 5.5 · Claude

Gaps are described as specific behaviors (e.g., not testing before building, saying yes to single-customer requests) with what good looks like.

Gemini 3.5 Flash-Lite · Gemini

Gaps are described as relying on outbound requests without validation, and what good looks like is interrogating the why and obsessing over impact.

Weighs the whole periodRightRightRight
GPT-6 Luna · API

The review covers the full year, referencing Q3 1:1, full-year metrics, and the entire set of shipped features, without over-weighting any single event.

Opus 5.5 · Claude

The review covers the full year, explicitly references the Q3 1:1, and does not let one recent event dominate.

Gemini 3.5 Flash-Lite · Gemini

Covers the full year, references Q3 1:1, and doesn't let one event dominate.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.8 FlashwithAPI83.357.72None
7Gemini 3.5 Flash-LitewithGemini73.258.63None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review