Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty66% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly71% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM working for Hannah Iyer, Quarry's CPO. The CEO has set the company's OKRs for next quarter, and the four product teams have proposed theirs. Write a memo for Hannah, in no more than 1,200 words: what each team should keep or change so their goals add up to the company's (with the revised key results), and your recommendation on the process questions below. The pack below is everything we have. Not all of it matters equally.

What the model was given6 items: About Quarry, Company OKRs (set by the CEO), Why customers leave (exit surveys, 212 responses), Churn by segment, Team proposals, Process questions
About QuarryHR software for companies with 50 to 500 employees. 420 staff, 2,600 customers. The core HR product is sold per employee; Payroll is a paid add-on.
Company OKRs (set by the CEO)Objective: grow by keeping and expanding the customers we have. Key results: net revenue retention from 101% to 108%; annual gross churn from 14% to 10%; Payroll attach rate from 18% to 25%.
Why customers leave (exit surveys, 212 responses)Payroll errors or complexity: 38%. Price: 24%. Missing integrations: 19%. Moved to an all-in-one competitor: 11%. Other: 8%.
Churn by segmentCustomers using Payroll: 8% a year. Customers without it: 16%. Customers with three or more integrations connected: 6%.
Team proposalsCore HR. Objective: build the best HR platform. Key results: ship the org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. Payroll. Objective: make payroll effortless. Key results: cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. Integrations. Objective: connect everything. Key result: ship 12 new integrations. Growth. Objective: win more customers. Key results: new logos up 20%; trial conversion from 9% to 12%.
Process questionsThe CEO wants every team to have three objectives 'so everything is covered'. The CFO wants OKR scores to set PM bonuses from next quarter. Last year teams scored 0.95 on average. Planning last quarter took five of the quarter's thirteen weeks.
What a strong answer doesThe answer key the graders mark against

Rewrites each team's goals so they add up to the company's. Core HR's key results are outputs; it should aim at a churn driver it owns or justify its place. Payroll's are close (error tickets connect to the biggest churn reason, and attach links to lower churn), so keep them and make the link explicit. Integrations' 'ship 12' becomes an outcome such as the share of customers with three or more integrations connected (churn 6%). Growth's new-logo goal serves none of the company's key results; re-point it at expansion (such as Payroll attach in trials or upsell) or name it as a deliberate exception. Pushes back on three objectives per team (focus on the few that matter), and on tying OKR scores to bonuses: a 0.95 average already suggests safe targets, and pay would make it worse. Suggests a lighter, shorter planning process.

Critical failures (cap the score at 40)
  • Endorses tying OKR scores to PM bonuses without flagging the risk
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The memo is clear, actionable, and covers all required elements; a PM could use it with only minor wording adjustments.

Output

To: Hannah Iyer, CPO Subject: Next-quarter OKRs: focus teams on retention and expansion

The CEO’s objective is clear, but the team proposals do not yet add up to it: Growth is optimizing new logos, while Core HR and Integrations mostly measure shipping. I recommend one objective per team, with a small number of measurable results tied to the company outcomes. Keep delivery plans such as “ship 12 integrations” as initiatives, not proof of impact.

Team recommendations

Core HR — change the objective and replace output KRs. “Build the best HR platform” is too broad. Org-chart redesign, custom fields, and CSAT may be useful, but the pack does not show that shipping them will reduce churn or grow revenue. Prioritize work against the customer problems behind churn—including price/value, which needs a product and pricing response rather than a feature assumption.

Revised objective: Improve retention of core HR customers. - Reduce annualized gross churn for customers without Payroll from 16% to at most 10.7%. This is a first-pass allocation: if Payroll attach reaches 25%, Payroll churn stays at 8%, and the segment rates are customer-weighted, non-Payroll churn needs to be about 10.7% for overall churn to reach 10%. Validate definitions and denominators before locking this target. - Keep CSAT at or above 4.2 as a guardrail, rather than making a score increase the main success measure.

Treat the two proposed launches as candidate initiatives, to be prioritized only if evidence links them to retention or expansion.

Payroll — keep the strongest proposal; clarify the outcome. Payroll directly supports both the company’s attach and churn goals. Payroll customers churn at half the rate of non-Payroll customers, and payroll errors or complexity are the most common reported reason for leaving.

Revised objective: Make Payroll reliable and grow adoption among current customers. - Increase Payroll attach from 18% to 25%. - Reduce payroll-error tickets by 40%. - Keep annual Payroll customer churn at 8% or lower.

Define error-ticket counting consistently, and pair the reduction with a quality guardrail so the metric cannot improve by discouraging customers from reporting problems.

Integrations — keep the problem area; replace “ship 12” as the key result. The evidence supports this area: missing integrations account for 19% of exit-survey responses, and customers with three or more integrations have 6% annual churn. But shipping integrations does not show that customers adopt them or stay.

Revised objective: Reduce integration-related customer loss. - Reduce integrations’ share of exit-survey reasons for leaving from 19% to 14%; treat this as a directional diagnostic, since survey shares are noisy. - Increase the number of customers with three or more active integrations. Establish the baseline in week one and set a numeric quarter target before plans are finalized; the pack does not provide enough data to invent one. - Keep “12 integrations shipped” as a delivery commitment only where customer demand and expected adoption justify it.

Growth — change the objective and stop optimizing acquisition this quarter. New logos and trial conversion do not support the stated company objective. Reorient Growth toward expansion in the installed base, working with Payroll on customer targeting and adoption.

Revised objective: Grow revenue from existing customers. - Increase company NRR from 101% to 108%; Growth is accountable for the expansion motion, with other teams contributing to the shared result. - Set a numeric target, after establishing a baseline, for expansion revenue from existing customers excluding Payroll attach. This makes Growth’s contribution visible without double-counting Payroll’s attach KR.

Process recommendations

  • Do not require three objectives per team. Coverage is not a reason to dilute focus. Use one objective and a few KRs per team; capture other responsibilities as initiatives or guardrails. Make cross-team ownership explicit, with one directly accountable owner for each company KR.
  • Do not use OKR scores to set PM bonuses next quarter. Bonuses would encourage teams to choose safe targets, avoid shared work, or optimize the score instead of customer outcomes. An average score of 0.95 last year is already a reason to examine target-setting and scoring calibration—not to add financial stakes. Keep performance and compensation decisions separate from quarterly OKR scores.
  • Cut planning from five weeks to two. Five weeks consumes nearly 40% of the quarter. Timebox baseline checks and target-setting, then revisit assumptions during the quarter. Use the missing integration and expansion baselines to sharpen targets early, not to prolong planning.

Finally, treat the company’s annual churn and NRR measures as lagging outcomes: review leading indicators during the quarter, but keep the company outcomes as the test of whether the portfolio is working.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision68%
  • passRespects explicit constraints54%
  • passIdentifies material uncertainty65%
  • passAvoids unsupported claims30%
  • partialProduces the required deliverable52%
  • passAims the teams with the churn data94%
  • passReads the 0.95 average for what it is91%
  • passKey results are outcomes, not output70%
  • passFocuses on the big rock97%
  • passShows how the goals add up57%
Run
Run
#1
API response time
57 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 11 checks passed

The memo is complete, well-structured, and actionable; a PM could use it with minor edits to fill in baselines, which the memo already flags as needed.

Output

MEMO To: Hannah Iyer, CPO From: Staff PM Re: Aligning team OKRs to the company OKRs, and the process questions

Bottom line

Only Payroll's proposal is close to the company objective (keep and expand existing customers). Core HR and Integrations propose output goals (features and integrations shipped). Growth targets new logos, which move none of the CEO's three key results. I recommend one objective per team, no mechanical link between OKR scores and bonuses, and planning capped at two weeks.

The logic behind the changes

Applying the exit-survey shares to 14% annual churn gives roughly:

  • Payroll errors/complexity: 5.3 points
  • Price: 3.4 points
  • Missing integrations: 2.7 points
  • All-in-one competitor: 1.5 points
  • Other: 1.1 points

Getting to 10% means removing about 4 points.

Payroll attach alone won't get there. Moving attach from 18% to 25% is about 180 more customers. Payroll customers churn at 8% against 16% for the rest, so even if the gap were entirely causal, attach is worth about 0.6 points of churn. Much of the gap is probably selection, because engaged customers adopt more. The same caution applies to the 6% churn among customers with three or more integrations.

Attach could also backfire. Payroll errors are the top exit reason, so pushing a flawed product onto more customers risks creating more of the churn we are trying to prevent.

The bulk of the churn reduction has to come from fixing Payroll quality and integration gaps. I suggest this split, as a planning assumption rather than a forecast:

Exit poolOwnerTargetChurn points
Payroll errors/complexityPayrollcut ~40%~2.1
Missing integrationsIntegrationscut ~33%~0.9
PriceGrowthcut ~15%~0.5
All-in-one competitorCore HRcut ~33%~0.5

Together these come to about 4 points, which is 14% to 10%.

The exit data has limits. We have 212 responses against roughly 360 annual churners, and "payroll complexity" may include non-users who left for a payroll-capable competitor. Week one should confirm the cuts by segment.

Team by team

Payroll: keep most of it, tighten it. - Objective: Make payroll reliable enough that customers stay and others add it. - KR1: Payroll-error tickets per 1,000 pay runs down 40%. Normalise the metric. A raw ticket count will rise as attach grows, and it can also be gamed by making tickets harder to file. - KR2: Payroll-error/complexity-cited exits down about 40% (about 2 points of company churn). - KR3: Payroll attach from 18% to 25%, with a guardrail that churn among Payroll customers stays at or below 8%. Sequence the attach push so it leans on the error fixes, targeting segments with clean data first.

Integrations: change the unit of success. "Ship 12" measures output, and the number is arbitrary. The 19% of exits citing missing integrations tells us which ones matter. - Objective: Make Quarry the system customers can't easily unplug. - KR1: Share of customers with three or more integrations connected, up from [baseline needed] by an agreed number of points. - KR2: Integration-cited exits down about a third (about 0.9 points of churn). - Ship the integrations ranked by exit and lost-customer demand. That list is the initiative that serves the KRs, not a KR itself.

Growth: the biggest change. Point it at existing customers. New logos and trial conversion don't affect NRR, churn or attach. A 20% rise in new logos can even worsen churn if the new customers are poorly fit. - Objective: Grow revenue from customers we already have. - KR1: NRR from 101% to 108%. Growth owns the expansion motion, and the other teams contribute. - KR2: Price-cited exits down about 15% through packaging and bundling experiments (for example Core HR plus Payroll). Don't raise list prices this quarter, since price is already the second-largest exit reason. - KR3: Share of new customers with three or more integrations connected within 90 days (baseline to be set). This is a lifecycle problem, which is a growth skill. - Trial conversion (9%) stays on the dashboard with a floor, but it is not a KR. If you want new-logo growth kept as a goal, that is a conversation with the CEO, because it contradicts his objective.

Core HR: the weakest link to the company goal. "Best HR platform" can't be scored, and two of the three KRs are ship dates. I have no evidence that the org-chart redesign or custom fields v2 affect retention. They may be worth building, but they don't earn a KR slot without that link. - Objective: Make Core HR the reason customers don't need an all-in-one competitor. - KR1: All-in-one-cited exits down about a third (about 0.5 points of churn). - KR2: Share of customers whose employee data is complete enough to switch on Payroll, up from [baseline needed]. This makes Core HR a direct contributor to attach, and Core HR and Payroll should confirm the definition. - KR3: CSAT from 4.2 to 4.5, kept as a health metric. - Org chart and custom fields stay on the roadmap only if exit or customer data shows they close the competitor gap. Core HR should make that case in week one.

Process recommendations

1. Three objectives per team: decline, with a better answer to "covered." The company has one objective with three KRs. Four teams with three objectives each would create twelve objectives pulling in different directions, which is the problem this memo is fixing. It would also lengthen planning. I recommend one objective and two or three KRs per team. To address the CEO's concern about coverage, give each team a short health-metric list with floors, such as uptime, support SLAs and CSAT. Those get monitored without being goals. A team may request a second objective if it can show a distinct customer outcome that the company objective doesn't cover.

2. OKR scores setting PM bonuses: recommend against. Last year's average score of 0.95 means the goals weren't stretching or the scoring was generous. Either way, the scores carry little information. Tying pay to them would make this worse, because teams would sandbag, avoid ambitious outcomes and game the metrics (ticket counts being the obvious example). The revised KRs are outcome-based and only partly within a team's control, which makes mechanical payouts less fair still.

Instead: - Calibrate so a typical score is around 0.7, with explicit committed and stretch levels. - Use scores as one input to performance reviews, judged by the manager in context. - Revisit a compensation link after two quarters of calibrated scoring.

If the CFO needs something concrete, tie any variable pay to the three company-level KRs, not team scores. That gives a shared incentive to cooperate.

3. Planning took five of thirteen weeks (38% of the quarter): cap it at two. - The CEO fixes company OKRs at least two weeks before the quarter starts, so teams never plan in the dark. - Teams draft on a one-page template with one objective and no more than three KRs, each with a baseline, owner and target. - One 90-minute cross-team session reconciles overlaps, using the exit-reason allocation above as the starting point. - Planning overlaps the last two weeks of the prior quarter and ends with sign-off by day 5 of the new one. - Fewer objectives per team will itself shorten the cycle.

Decisions and next steps

  1. Approve the reframing in this memo, and take the Growth change to the CEO before teams redraft, since it is the one most likely to be contested.
  2. Ask Data for baselines this week: share of customers with three or more integrations, exit reasons by segment, and payroll tickets per pay run. Several KRs above can't be finalised without them.
  3. Ask each team to resubmit within one week.
  4. Take the bonus and three-objective positions to the CEO and CFO together, so we have one conversation rather than two.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision85%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims52%
  • partialProduces the required deliverable66%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is99%
  • passKey results are outcomes, not output60%
  • passFocuses on the big rock99%
  • passShows how the goals add up60%
Run
Run
#1
API response time
67 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 11 checks passed

The memo is complete, evidence-based, and actionable; a PM could use it with only minor wording adjustments, without re-checking facts or redoing analysis.

Output

To: Hannah Iyer, CPO From: Staff PM Re: Q-next OKRs: making team goals add up, and the process questions

The short version

As proposed, the four teams' OKRs do not add up to the company's. Only Payroll's goals clearly support the CEO's objective. Integrations is measuring output rather than outcome. Core HR is pointed at its own roadmap. Growth is pointed at new customers, which the company OKRs deliberately leave out.

I recommend one objective per team, each tied to one or more of the three company KRs, with outcome-based key results. On process, I recommend no to three objectives per team, no to bonuses tied to OKR scores in their current form, and a two-week cap on planning.

How the company numbers can be reached

The company KRs are linked. Net revenue retention (NRR) is roughly 100% minus churned revenue plus net expansion.

  • Churn. Cutting gross churn from 14% to 10% is worth about 4 points of NRR.
  • Expansion. The other ~3 points have to come from expansion. Because Payroll is a paid add-on, attach (18% → 25%) is the main lever we control.
  • What finance should check. Whether 7 points of attach yields ~3 points of NRR depends on Payroll's price relative to core seats. That is worth checking before we commit to the 108% target.

Where churn comes from:

  • Exit reasons (212 surveys, about 364 churned customers a year): payroll errors or complexity 38%, price 24%, missing integrations 19%, moved to an all-in-one competitor 11%.
  • Payroll users churn at 8%, versus 16% for customers without Payroll.
  • Customers with three or more integrations churn at 6%.

Two caveats shape the recommendations:

  1. Attach alone barely moves churn. Moving attach from 18% to 25% at today's segment churn rates lowers blended churn by only about half a point (from roughly 14.6% to 14% on these figures). And that assumes the relationship is causal; engaged customers may simply buy Payroll and stay. Most of the 4-point churn reduction has to come from fixing payroll errors and from integrations.
  2. Payroll errors are the biggest single churn reason. They must be fixed before we push attach hard. Selling more Payroll while it still drives 38% of exits could raise churn rather than reduce it.

Team by team

Payroll: keep, with small changes - Objective: keep "Make payroll effortless". - Keep: cut payroll-error tickets by 40%, as a leading indicator. - Add: gross churn among Payroll customers from 8% to 6%. This is the outcome that matters, and tickets can fall without customers noticing. - Change ownership of attach: the 18% → 25% attach KR should be jointly owned with Growth. Payroll can make the product sellable; it does not run the cross-sell motion. - Sequencing: error fixes ship in the first half of the quarter, and the attach push ramps after.

Integrations: change the key result - Objective: change to "Make Quarry the hub customers can't unplug." - Replace "ship 12 integrations" with: share of customers with three or more integrations connected, from [baseline] to [baseline + 10 pts]. The baseline should be confirmed in week one. - Add: activation rate of new integrations within 60 days of launch. - How to choose integrations: prioritise by the systems named in exit surveys and churn-risk accounts, not by count. Four integrations customers actually connect beat twelve they don't. - Test the causal claim: run a controlled push to drive integration adoption in one cohort. The 6% churn figure may reflect selection rather than effect.

Core HR: refocus - Problem: org-chart redesign and custom fields v2 are deliverables, not results. "Build the best HR platform" has no line of sight to retention. CSAT is not a company KR. - Objective: change to "Give customers without Payroll a reason to stay." This is the 16%-churn segment, which is most of the base. - New KR 1: gross churn among non-Payroll customers from 16% to 13%. - New KR 2: a usage-based health score for at-risk accounts, with a target set once the baseline is known. - Existing projects: keep org chart and custom fields only if the team can show they address retention. For example, custom fields v2 could matter if integrations need them for data mapping. - Capacity: if the team can't show that link, I'd lend some of its engineering capacity to Payroll's error work for the quarter.

Growth: redirect - Problem: new logos and trial conversion don't count toward NRR, churn or attach. New customers are excluded from NRR by definition. - Objective: change to "Expand within our existing customers." - New KR 1: Payroll attach 18% → 25%, shared with Payroll, through in-product trials and cross-sell for existing customers. - New KR 2: conversion from Payroll trial to paid among existing customers, from [baseline] to [target]. - Your call before we tell Growth: if the CEO still wants new-logo growth, that is a fourth company KR and should be added explicitly. Otherwise Growth is working off-strategy.

A gap no product team owns

Price drives 24% of exits, and no team can own that through product work alone. Price is also often a polite stand-in for "not worth it", which the work above partly addresses. Packaging and pricing need an owner from commercial or finance. I suggest you raise this with the CEO and CFO rather than assign it to a product team.

Revised OKRs at a glance

TeamObjectiveKey resultsCompany KR served
PayrollMake payroll effortlessError tickets −40%; Payroll-customer churn 8% → 6%; attach 18% → 25% (shared)Churn, attach, NRR
IntegrationsMake Quarry the hub customers can't unplugCustomers with 3+ integrations +10 pts; 60-day activation of new integrationsChurn, NRR
Core HRGive non-Payroll customers a reason to stayNon-Payroll churn 16% → 13%; health-score targetChurn, NRR
GrowthExpand within our existing customersAttach 18% → 25% (shared); existing-customer Payroll trial → paidAttach, NRR

Process recommendations

1. Three objectives per team: recommend against. - The company itself has one objective. Twelve team objectives would spread effort thin and reproduce the problem this memo fixes: goals that cover everything and add up to nothing. - Proposal for the CEO: one objective per team, two or three KRs each, all tied to company KRs. - Meeting his real concern: "everything covered" belongs in a health-metrics dashboard (uptime, support SLAs, CSAT) that each team maintains. That separates running the business from changing it, and gives him the coverage he wants without diluting focus.

2. OKR scores setting PM bonuses: recommend against as proposed. - Last year's average score of 0.95 already suggests goals were set to be hit, not stretched. A healthy stretch average is roughly 0.6–0.7. - Paying on scores will push every team to sandbag, just as the company sets genuinely hard targets like churn from 14% to 10%. It also discourages the honest scoring we need to learn what works. - Alternatives for the CFO: - Tie a portion of PM bonuses to company KR results, shared across teams, which encourages collaboration on the shared attach and churn goals. - Base the individual portion on manager assessment of the quality of the bets and the execution, with OKR scores as one input rather than a formula. - If the CFO insists on a formula: at minimum, calibrate targets centrally so teams can't set their own bar.

3. Planning took 5 of 13 weeks: cap it at two, finished before the quarter starts. - Week −3: company OKRs final. This quarter's are already set. - Week −2: teams draft one objective and two or three KRs against them. This memo can be the template. - Week −1: one joint review with you to resolve shared ownership and gaps. Then lock. - During the quarter: baselines marked [TBD] get filled in during week one. They should not reopen the plan.

Decisions I need from you

  1. Agree the revised team OKRs above, especially redirecting Growth.
  2. Decide whether new-logo growth is in or out. If in, it goes to the CEO as an explicit company KR.
  3. Raise pricing ownership, and the process recommendations, with the CEO and CFO.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly9%
  • passAddresses the actual decision89%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims35%
  • partialProduces the required deliverable60%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is100%
  • passKey results are outcomes, not output31%
  • passFocuses on the big rock99%
  • passShows how the goals add up85%
Run
Run
#1
Time to output
54 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 1

Uses the supplied evidence correctlyRightMixedMixed
GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context or simple arithmetic.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the supplied context or derived by arithmetic from it.

Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic; no invented numbers or data.

All got right 10

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

The memo commits to clear revised key results for each team and explicit process recommendations, framed for Hannah.

Sonnet 5.5 · API

The memo commits to a clear set of revised team OKRs and process recommendations, and states what would change the call (e.g., CEO conversation on new logos).

Opus 5.5 · Claude

The memo commits to clear revised OKRs and process recommendations, framed for the CPO, and states what would change the call (e.g., CEO's stance on new logos).

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output is a memo under 1,200 words, addresses the CPO, and respects the requested form.

Sonnet 5.5 · API

The output is a memo to Hannah, within the 1,200-word limit, and addresses the requested decisions.

Opus 5.5 · Claude

The output is a memo addressed to Hannah Iyer, stays within the 1,200-word limit, and respects the requested form.

Identifies material uncertaintyRightRightRight
GPT-6 Luna · API

It names missing baselines and definitional unknowns, and says to establish baselines in week one and validate definitions before locking targets.

Sonnet 5.5 · API

The memo identifies limits in the exit data, selection bias in churn segments, missing baselines, and says how to resolve them in week one.

Opus 5.5 · Claude

It names key unknowns (causal vs. selection effects for integrations and Payroll, NRR arithmetic, baselines) and suggests how to resolve them.

Avoids unsupported claimsRightRightRight
GPT-6 Luna · API

Interpretations and forecasts are clearly labelled or are directly supported by the evidence; no unsupported factual claims are presented as established.

Sonnet 5.5 · API

Hypotheses like selection bias and risks are clearly labelled as such; no interpretations are presented as established fact.

Opus 5.5 · Claude

Hypotheses and causal claims are clearly labelled as such (e.g., 'assumes the relationship is causal', 'may reflect selection rather than effect').

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The memo provides revised team goals with key results and process recommendations, and is usable with light edits.

Sonnet 5.5 · API

The memo provides complete revised KRs and process recommendations that Hannah could act on with light edits.

Opus 5.5 · Claude

The memo provides complete revised OKRs, process recommendations, and a clear ask; the CPO could act on it with minimal edits.

Aims the teams with the churn dataRightRightRight
GPT-6 Luna · API

It uses the churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll churn, with explicit links.

Sonnet 5.5 · API

The memo uses exit survey shares and churn-by-segment data to aim Payroll at errors, Integrations at 3+ integrations, Growth at price, and Core HR at all-in-one competitor.

Opus 5.5 · Claude

It uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll customers, with explicit links.

Reads the 0.95 average for what it isRightRightRight
GPT-6 Luna · API

It flags the 0.95 average as a sign of safe targets and recommends against tying bonuses to scores to avoid further sandbagging.

Sonnet 5.5 · API

The memo flags the 0.95 average as evidence of safe targets, warns tying bonuses would worsen sandbagging, and recommends against it.

Opus 5.5 · Claude

It flags that the 0.95 average suggests safe targets and argues that tying bonuses to scores would worsen sandbagging, with a clear recommendation against.

Key results are outcomes, not outputRightRightRight
GPT-6 Luna · API

Every key result is a measurable outcome (churn, attach, error tickets, CSAT guardrail, exit-survey share, 3+ integrations, NRR, expansion revenue); shipping appears only as initiatives.

Sonnet 5.5 · API

Every revised key result is a measurable outcome (e.g., error tickets per pay run, exit reductions, attach rate) with baselines and targets; shipping work is explicitly kept as initiatives.

Opus 5.5 · Claude

Every revised key result is a measurable outcome (churn, attach, activation, health score) with baselines and targets; shipping work is treated as initiatives.

Focuses on the big rockRightRightRight
GPT-6 Luna · API

It cuts each team to one objective with a few KRs, drops shipping KRs, and explains what was dropped and why.

Sonnet 5.5 · API

The memo cuts each team to one objective with 2-3 KRs, drops new logos and shipping goals, and explains why they don't serve the company objective.

Opus 5.5 · Claude

It cuts each team to one objective with a few KRs, explicitly drops non-contributing work (org chart, custom fields, new logos, shipping count), and explains why.

Shows how the goals add upRightRightRight
GPT-6 Luna · API

Each team-level KR is linked to a company goal with reasoning or data, and the new-logo goal is flagged as not supporting the company objective and reoriented.

Sonnet 5.5 · API

Every team KR is linked to a company KR (churn, NRR, attach) with arithmetic or reasoning, and the new-logo goal is flagged as serving none and recommended for removal.

Opus 5.5 · Claude

The table and narrative link every team KR to a company KR (churn, attach, NRR), flag Growth's original new-logo goal as off-strategy, and note the unowned pricing gap.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review