Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty66% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly71% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Pathwise, working for the CPO, Grace Mensah. Last quarter's OKR scorecard is below, along with the company goal it was meant to serve. Grace wants a memo of no more than 1,100 words: what last quarter's OKRs actually tell us (not just the scores), and next quarter's OKRs for the Learner app, Admin and Growth teams, each with baselines and targets, that add up to the company goal.\n\nThe files are attached.

What the model was given5 items: About Pathwise, company_okr_q3.md, okr_scorecard_q3.csv (team_label is the type each team gave its key result), metric_dictionary.md, churn_reasons_q3.md
About PathwiseOnline learning for companies: courses employees take, and tools for the HR and L&D teams who buy it. Sold per seat on annual contracts.
company_okr_q3.md7 lines · Download
# Company OKR, Q3 (set by the CEO)

Objective: grow by keeping and expanding the customers we have.
- Net revenue retention: from 103% to 108%. Actual: 102%.
- Quarterly logo churn: from 3.1% to 2.5%. Actual: 3.4%.

Q3 average key result score across product teams: 0.81. The all-hands slide said "a great quarter".
okr_scorecard_q3.csv (team_label is the type each team gave its key result)team,objective,key_result,team_label,baseline,target,actual,score Content,Best course library in L&D,Publish 40 new courses,output,0,40,41,1.00 Content,Best course library in L&D,Course completion rate,outcome,41,55,58,1.00 Content,Best course library in L&D,Translate top 20 courses into 3 languages,output,0,60,60,1.00 Learner app,Learning people love,Ship mobile offline mode,output,0,1,1,1.00 Learner app,Learning people love,Weekly active learners per seat,outcome,0.31,0.4,0.29,0.00 Learner app,Learning people love,Ship streaks and badges,output,0,1,1,1.00 Learner app,Learning people love,App store rating,outcome,4.2,4.6,4.5,0.75 Admin,Admins get value fast,Launch 3 HRIS integrations,outcome,0,3,3,1.00 Admin,Admins get value fast,Days from contract to first learner,outcome,34,14,31,0.15 Admin,Admins get value fast,Ship new reporting dashboard,output,0,1,1,1.00 Growth,Expand inside accounts,Seats added in existing accounts,outcome,9200,14000,8100,0.00 Growth,Expand inside accounts,Run 12 expansion experiments,output,0,12,13,1.00 Growth,Expand inside accounts,Launch in-app upgrade prompts,output,0,1,1,1.00 Platform,Reliable at scale,Uptime,outcome,99.5,99.9,99.93,1.00 Platform,Reliable at scale,Migrate to new video CDN,output,0,1,1,1.00 Platform,Reliable at scale,Close 150 bugs,output,0,150,162,1.00
metric_dictionary.md5 lines · Download
# Metric dictionary (excerpt)

- Course completion rate: share of course enrolments marked complete. Until 31 July "complete" meant every module finished. From 1 August it means 80% of modules finished (changed so completion matches the certificate rule).
- Weekly active learners per seat: learners who start or continue a course in a week, divided by paid seats.
- Days from contract to first learner: calendar days from contract signature to the first learner starting a course.
churn_reasons_q3.md8 lines · Download
# Why accounts left in Q3 (exit interviews, 38 accounts)

- "Our people didn't use it": 17
- Budget cut: 9
- Moved to the LMS that came with our HR system: 7
- Content didn't fit our industry: 5

…
What a strong answer doesThe answer key the graders mark against

Reads the scorecard for what it is: the average is 0.81, but every output key result scored 1.00 (things shipped) while outcome key results averaged about 0.48, and the company missed both goals (NRR 102% against 108%, churn 3.4% against 2.5%). Two scores flatter: 'Launch 3 HRIS integrations' is labelled an outcome but is output, and course completion only 'beat' its target because completion was redefined on 1 August (all modules to 80%), so the result can't be compared with its baseline. The outcome key results that matter for churn moved backwards or barely (weekly active learners per seat 0.31 to 0.29; days to first learner 34 to 31; seats added 8,100 against 9,200). Builds next quarter's OKRs on the churn evidence: low usage is the top exit reason, and accounts reaching a first learner within 14 days churn at a third of the rate. So Admin owns days from contract to first learner (from 31), Learner app owns weekly active learners per seat (from 0.29), Growth owns seats added in existing accounts or net expansion, each with a baseline from the files and a target, with shipping work as initiatives. Notes the completion metric needs restating on one definition before anyone uses it.

Critical failures (cap the score at 40)
  • Describes last quarter as a success on the strength of the 0.81 average
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, data files, staff-level

The outputs

3 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit10 of 11 checks passed

The memo contains a factual error in the average score calculation (0.74 vs 0.81), which would require re-checking the evidence before use.

Output

To: Grace Mensah, CPO From: Staff PM, Product Re: What Q3's OKRs tell us, and proposed Q4 OKRs for Learner app, Admin and Growth

---

1. Bottom line

Q3 was not a great quarter. Both company key results moved backwards:

  • Net revenue retention (NRR) fell from 103% to 102%, against a target of 108%.
  • Logo churn rose from 3.1% to 3.4%, against a target of 2.5%.

The 0.81 average hid this because it mixed outputs, which we control, with outcomes, which customers control. We shipped everything we said we would. The customer behaviours those launches were meant to change didn't move, or got worse.

2. What Q3 actually tells us

Outputs scored perfectly. Outcomes didn't.

  • All 10 output key results (KRs) scored 1.00. I'm counting "Launch 3 HRIS integrations" as an output; it was labelled an outcome, but it's a launch.
  • The 6 genuine outcomes averaged 0.48.
  • The three outcomes that link most directly to the company goal averaged 0.05: weekly active learners per seat (0.00), days to first learner (0.15) and seats added in existing accounts (0.00).

The 0.81 doesn't reproduce. The 16 KRs on the scorecard average 0.74. Averaging by team gives 0.75. We should find out how 0.81 was calculated before it's quoted again.

Launches didn't move the metrics they targeted.

TeamWhat shippedWhat happened
Learner appOffline mode, streaks and badgesWeekly active learners per seat fell from 0.31 to 0.29
Growth13 experiments, in-app upgrade promptsSeats added fell from 9,200 to 8,100
Admin3 HRIS integrations, new dashboardDays to first learner improved only from 34 to 31, against a target of 14

One "win" is probably a definition change. Course completion rose from 41% to 58%. But on 1 August, "complete" changed from 100% of modules finished to 80%. Completion climbing while weekly activity fell suggests most of the gain came from the new definition. We should restate Q3 on both definitions before claiming it.

Scoring rewarded going backwards no more harshly than standing still. Weekly activity and seats added both ended below their baselines, yet each scored 0.00, the same as no change. The scorecard couldn't show that we lost ground.

The churn data points to adoption. Of 38 exit interviews:

Reason for leavingAccountsShare
"Our people didn't use it"1745%
Budget cut924%
Moved to the LMS bundled with their HR system718%
Content didn't fit their industry513%

Accounts that reached a first learner within 14 days churned at one third the rate of the rest. That is a correlation, not proof: healthier buyers may simply onboard faster. Still, it is our strongest signal. Adoption is the lever. More features and more experiments are not, on Q3's evidence.

3. Rule changes for Q4

  1. KRs are outcomes only. Launches go in a separate initiatives list and are not scored.
  2. Score against the baseline, and allow negative scores. Ending below the baseline should look worse than standing still.
  3. Freeze metric definitions for the quarter. Any change gets restated on both the old and new definition.
  4. Every KR names the company KR it serves.

4. Proposed Q4 OKRs

Baselines are Q3 actuals where we have them. Three baselines are marked TBC: we don't yet track those metrics, and Data will deliver them by the end of week 2. Each target is set as a fixed increase on its baseline, so the target is fixed as soon as the baseline is known.

Learner app. Objective: the seats customers pay for get used. This serves both churn and NRR.

KRBaselineTarget
Weekly active learners per paid seat (all accounts)0.290.33
Weekly active learners per seat in accounts renewing in Q4 or Q1TBCBaseline + 0.05
Guardrail: app store rating4.5Stays at or above 4.5

The second KR matters because it targets the accounts whose renewal decision is being made now.

Admin. Objective: every new customer has learners in their first two weeks. This serves churn.

KRBaselineTarget
Median days from contract to first learner31 (Q3 figure; confirm it is a median, not a mean)20
Share of new accounts with a first learner within 14 daysTBCBaseline + 25 points
Share of new accounts with an HRIS sync live by day 14TBC50%

The HRIS sync KR turns the integrations we shipped into usage. It also addresses the 7 accounts that left for their HR system's bundled LMS.

Growth. Objective: expand where usage proves value, and stop seat losses at renewal. This serves NRR.

KRBaselineTarget
Seats added in existing accounts8,10010,000
Seat retention at renewal (seats renewed ÷ seats up for renewal)TBC (Finance)Baseline + 3 points

Growth should focus expansion on accounts above a usage threshold, to be agreed with Learner app. Q3 showed that prompting low-usage accounts to upgrade doesn't work.

5. How these add up to the company goal, and where they don't

Churn (3.4% → 2.5%). This works through two levers:

  • Q4–Q1 renewal accounts. Contracts are annual, so churn this quarter is mostly decided by how accounts used us over the past year. The only lever with a near-term effect is raising usage in accounts renewing soon. That is the Learner app's second KR.
  • New-account onboarding. Admin's onboarding KRs act on new accounts, which won't come up for renewal for a year.

To illustrate the size of the onboarding effect: suppose 20% of accounts currently reach a first learner within 14 days, and those accounts churn at one third the rate of the rest. Moving that share to 50% would cut churn from 3.4% to about 2.6%. That only holds once the whole customer base has passed through the new onboarding, and only if the relationship is causal.

Honest read: we will not reach 2.5% churn in Q4. Admin's onboarding work is what gets us there in 2025.

NRR (102% → 108%). NRR gains come from two places: fewer churned and contracted seats (Growth's seat retention KR, plus the Learner app's usage work) and more seats added (Growth's expansion KR).

I can't yet show the arithmetic that closes the 6-point gap. That needs Finance's figures for revenue up for renewal and average price per seat. Even with the 10,000 seats target, I expect Q4 to land nearer 104–105% than 108%.

The remaining 6 churned accounts don't map to these KRs, though the HRIS work partly addresses the bundled-LMS losses:

  • Budget cuts (9 accounts) need commercial responses, such as flexible tiers.
  • Content fit (5 accounts) belongs to the Content team. Their Q4 KR should be usage of newly published courses, not the number of courses published.

6. Decisions I need from you

  1. Approve the four scoring rules, and correct the 0.81 figure with the CEO.
  2. Discuss the company targets with the CEO. I suggest keeping 108% NRR and 2.5% churn as the year-end destination, with Q4 checkpoints of 104–105% NRR and 3.0% churn, so the scorecard tells the truth.
  3. Ask Finance and Data to deliver the TBC baselines and the NRR arithmetic by the end of week 2. I'll lock the targets the same day.

Claims the judge couldn’t find in the brief

  • The 16 KRs on the scorecard average 0.74, not 0.81.

Grades and run details

Decision model 95 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly28%
  • passAddresses the actual decision82%
  • passRespects explicit constraints17%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims47%
  • passProduces the required deliverable39%
  • passReads the scorecard for what it is100%
  • passCatches the redefined key result100%
  • passKey results are outcomes, not output79%
  • passFocuses on the big rock67%
  • partialShows how the goals add up25%
Run
Run
#1
Time to output
65 s
Submitted
5 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Mixed 1

Uses the supplied evidence correctlyMixed

The claim that the 16 KRs average 0.74 is a factual error; the correct average is 0.81, so not all statements about the current situation are supported.

Got right 10

Addresses the actual decisionRight

Commits to proposed Q4 OKRs and says what would change the NRR target (Finance figures).

Respects explicit constraintsRight

Memo format, under 1,100 words, includes Q3 analysis and Q4 OKRs for the three teams with baselines and targets.

Identifies material uncertaintyRight

Identifies TBC baselines, correlation vs causation of onboarding, and NRR gap needing Finance data, with resolution steps.

Avoids unsupported claimsRight

Interpretations are labelled as hypotheses (e.g., 'probably a definition change', 'correlation, not proof').

Produces the required deliverableRight

Complete memo to Grace, within length, with actionable Q4 OKRs and clear asks.

Reads the scorecard for what it isRight

Splits output (1.00) vs outcome (0.48) scores and contrasts with missed company goals.

Catches the redefined key resultRight

Catches the course completion redefinition and the integrations KR being output labelled as outcome, and explains the implications.

Key results are outcomes, not outputRight

All proposed KRs are measurable outcomes with baselines and targets; shipping work is explicitly separated.

Focuses on the big rockRight

Few objectives per team, explains what was dropped (e.g., low-usage upgrade prompts, budget cuts and content fit handled elsewhere).

Shows how the goals add upRight

Each KR is linked to churn or NRR, and uncovered churn reasons are flagged with suggested owners.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review