Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 20 graded outputs by 7 models. 65% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Pathwise, working for the CPO, Grace Mensah. Last quarter's OKR scorecard is below, along with the company goal it was meant to serve. Grace wants a memo of no more than 1,100 words: what last quarter's OKRs actually tell us (not just the scores), and next quarter's OKRs for the Learner app, Admin and Growth teams, each with baselines and targets, that add up to the company goal.\n\nThe files are attached.

What the model was given5 items: About Pathwise, company_okr_q3.md, okr_scorecard_q3.csv (team_label is the type each team gave its key result), metric_dictionary.md, churn_reasons_q3.md
About PathwiseOnline learning for companies: courses employees take, and tools for the HR and L&D teams who buy it. Sold per seat on annual contracts.
company_okr_q3.md7 lines · Download
# Company OKR, Q3 (set by the CEO)

Objective: grow by keeping and expanding the customers we have.
- Net revenue retention: from 103% to 108%. Actual: 102%.
- Quarterly logo churn: from 3.1% to 2.5%. Actual: 3.4%.

Q3 average key result score across product teams: 0.81. The all-hands slide said "a great quarter".
okr_scorecard_q3.csv (team_label is the type each team gave its key result)team,objective,key_result,team_label,baseline,target,actual,score Content,Best course library in L&D,Publish 40 new courses,output,0,40,41,1.00 Content,Best course library in L&D,Course completion rate,outcome,41,55,58,1.00 Content,Best course library in L&D,Translate top 20 courses into 3 languages,output,0,60,60,1.00 Learner app,Learning people love,Ship mobile offline mode,output,0,1,1,1.00 Learner app,Learning people love,Weekly active learners per seat,outcome,0.31,0.4,0.29,0.00 Learner app,Learning people love,Ship streaks and badges,output,0,1,1,1.00 Learner app,Learning people love,App store rating,outcome,4.2,4.6,4.5,0.75 Admin,Admins get value fast,Launch 3 HRIS integrations,outcome,0,3,3,1.00 Admin,Admins get value fast,Days from contract to first learner,outcome,34,14,31,0.15 Admin,Admins get value fast,Ship new reporting dashboard,output,0,1,1,1.00 Growth,Expand inside accounts,Seats added in existing accounts,outcome,9200,14000,8100,0.00 Growth,Expand inside accounts,Run 12 expansion experiments,output,0,12,13,1.00 Growth,Expand inside accounts,Launch in-app upgrade prompts,output,0,1,1,1.00 Platform,Reliable at scale,Uptime,outcome,99.5,99.9,99.93,1.00 Platform,Reliable at scale,Migrate to new video CDN,output,0,1,1,1.00 Platform,Reliable at scale,Close 150 bugs,output,0,150,162,1.00
metric_dictionary.md5 lines · Download
# Metric dictionary (excerpt)

- Course completion rate: share of course enrolments marked complete. Until 31 July "complete" meant every module finished. From 1 August it means 80% of modules finished (changed so completion matches the certificate rule).
- Weekly active learners per seat: learners who start or continue a course in a week, divided by paid seats.
- Days from contract to first learner: calendar days from contract signature to the first learner starting a course.
churn_reasons_q3.md8 lines · Download
# Why accounts left in Q3 (exit interviews, 38 accounts)

- "Our people didn't use it": 17
- Budget cut: 9
- Moved to the LMS that came with our HR system: 7
- Content didn't fit our industry: 5

…
What a strong answer doesThe answer key the graders mark against

Reads the scorecard for what it is: the average is 0.81, but every output key result scored 1.00 (things shipped) while outcome key results averaged about 0.48, and the company missed both goals (NRR 102% against 108%, churn 3.4% against 2.5%). Two scores flatter: 'Launch 3 HRIS integrations' is labelled an outcome but is output, and course completion only 'beat' its target because completion was redefined on 1 August (all modules to 80%), so the result can't be compared with its baseline. The outcome key results that matter for churn moved backwards or barely (weekly active learners per seat 0.31 to 0.29; days to first learner 34 to 31; seats added 8,100 against 9,200). Builds next quarter's OKRs on the churn evidence: low usage is the top exit reason, and accounts reaching a first learner within 14 days churn at a third of the rate. So Admin owns days from contract to first learner (from 31), Learner app owns weekly active learners per seat (from 0.29), Growth owns seats added in existing accounts or net expansion, each with a baseline from the files and a target, with shipping work as initiatives. Notes the completion metric needs restating on one definition before anyone uses it.

Critical failures (cap the score at 40)
  • Describes last quarter as a success on the strength of the 0.81 average
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, data files, staff-level

The outputs

6 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The memo is complete, evidence-based, and actionable; a PM could use it with minor wording edits without re-checking facts or redoing reasoning.

Output

MEMO To: Grace Mensah, CPO From: Staff PM Re: What Q3 OKRs tell us, and Q4 OKRs for Learner app, Admin and Growth

Bottom line

Q3 scored 0.81 and missed the company goal. NRR was 102% against a 103% baseline, and churn was 3.4% against 3.1%. Both got worse. Teams delivered what they planned, but the plan didn't move usage, and non-use is why customers leave. The Q4 OKRs below target usage and activation directly.

What Q3 actually tells us

1. The score measures delivery, not results. Ten of 16 KRs are outputs, and all ten scored 1.00. That count includes "Launch 3 HRIS integrations", which Admin labelled an outcome. The six true outcome KRs average 0.48. The 0.81 is mostly shipping.

2. What shipped didn't move the customer metrics. - Offline mode plus streaks and badges shipped, and weekly active learners per seat fell from 0.31 to 0.29. - Three integrations and a new dashboard shipped, and days to first learner moved only from 34 to 31. - Thirteen experiments and upgrade prompts shipped, and seats added fell from 9,200 to 8,100.

3. Customers told us why they left. Of 38 exits, 17 (45%) said their people didn't use it. Seven (18%) moved to their HRIS vendor's bundled LMS, nine cited budget and five cited content fit. Accounts reaching first learner within 14 days churned at a third of the rate of the rest. That could partly be selection, since healthier customers may simply move faster. It is still the strongest lead we have.

4. Two "wins" shouldn't be trusted. - Completion rate (41% to 58%) is not evidence. The definition loosened on 1 August from all modules to 80%, so the baseline and actual are measured differently. It needs restating on one definition before anyone cites it. - Uptime hit target, but no exit cited reliability. That is good hygiene, not a retention lever.

5. Process fix. Outputs should be initiatives, not KRs. Every team should carry at least one KR tied to a churn or NRR driver, and we should report outcome KRs separately from the average.

Q4 OKRs

Assumption (please confirm): the company goal carries forward unchanged: NRR 102% to 108% and quarterly logo churn 3.4% to 2.5%.

Learner app: Learners come back every week

Key resultBaselineQ4 target
Weekly active learners per seat0.290.35
Shared with Admin: churn attributed to non-use (pts of the 3.4)1.50.75
Guardrail: app store rating4.5≥4.5

Admin: Admins get learners learning in two weeks

Key resultBaselineQ4 target
Days from contract to first learner3121
Churn from accounts moving to an HRIS-bundled LMS (pts)0.60.45
Shared with Learner: churn attributed to non-use (pts)1.50.75

Growth: Expand where seats are used

Key resultBaselineQ4 target
Seats added in existing accounts8,10012,000

Growth's experiments should target accounts with healthy usage, since expansion into idle accounts is unlikely. Growth also depends on Learner's usage KR.

How it adds up

Churn: 3.4% to 2.5%. I split the 3.4 points by exit reason, assuming the 38 interviews are representative.

ReasonShare of exitsPts of churnQ4
Non-use45%1.50.75
Budget24%0.80.8
HRIS-bundled LMS18%0.60.45
Content fit13%0.40.4
Total3.4~2.4

Halving non-use churn takes 0.75 points, shared between Learner (usage) and Admin (activation). The HRIS reduction takes 0.15. That gets to about 2.4, leaving 0.1 of buffer. I've put nothing against budget or content fit, since neither is addressable this quarter.

NRR: 102% to 108%. I'm assuming a 0.9-point quarterly churn improvement is worth about 3.6 NRR points annualised. Expansion has to supply the other 2.4. The 12,000 seats target (about 30% above Q2's 9,200) is a placeholder sized to that gap. I don't have the seat base or ARR to convert it exactly.

Risks and asks

  • The churn target is a stretch. On annual contracts, Q4 churn comes from accounts renewing in Q4, so a 31-to-21-day onboarding improvement mostly pays off at later renewals. I want a list of accounts renewing before the end of Q1, with their usage. Learner and Admin should work that list first, so Q4 effort lands on accounts that can still change their decision. If we can't act on it, a realistic fallback is about 2.9% churn and 104-105% NRR. I'd rather agree that with the CEO now than explain it in January.
  • Finance: please confirm how churn converts to NRR and what 12,000 seats is worth in NRR points.
  • Data: restate completion rate on the 80% rule for Q2 and Q3.
  • Exit data: 38 interviews is a small sample, and I assumed it covers all churn. Please confirm.
  • Decision for you: confirm the Q4 company goal, and whether you'll back splitting the score into output and outcome in the next all-hands.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly10%
  • passAddresses the actual decision53%
  • partialRespects explicit constraints15%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims34%
  • partialProduces the required deliverable31%
  • passReads the scorecard for what it is93%
  • passCatches the redefined key result100%
  • passKey results are outcomes, not output81%
  • passFocuses on the big rock56%
  • passShows how the goals add up29%
Run
Run
#1
API response time
1.5 min
Submitted
6 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is well-structured, evidence-based, and provides actionable OKRs with clear reasoning; a PM could use it with minor wording tweaks and no rework.

Output

To: Grace Mensah, CPO From: Staff PM Subject: Q3 diagnosis and proposed Q4 OKRs

Recommendation

Q3 was a strong delivery quarter, not a strong customer-growth quarter. We should carry the missed company targets into Q4, reset baselines to Q3 actuals, and judge teams on customer behavior and retained revenue—not shipped features.

Proposed Q4 company objective: Grow by keeping and expanding existing customers. - Net revenue retention (NRR): 102% → 108%. - Quarterly logo churn: 3.4% → 2.5%.

These are recovery targets, not a forecast justified by last quarter’s results.

What Q3 actually tells us

The 0.81 score obscures the company’s deterioration. NRR fell from 103% to 102%; logo churn rose from 3.1% to 3.4%. After correctly classifying “launch three HRIS integrations” as an output, 10 of the 16 team KRs measured delivery. All ten scored 1.00. The six remaining measures averaged approximately 0.48—and one of those, completion, is not comparable over time. Equal-weight averaging rewarded delivery while masking misses on customer value.

Learner adoption is the clearest warning. Weekly active learners per paid seat fell from 31% to 29%, despite offline mode, streaks and badges shipping. App-store ratings improved from 4.2 to 4.5, but that does not establish broader adoption among employees whose employers pay for seats. “Our people didn’t use it” appeared in 17 of 38 exit interviews, the largest category. This supports prioritizing adoption, although exit interviews neither represent every customer nor establish causality.

Admin delivery barely changed time to value. Three integrations and a reporting dashboard shipped, but time from contract to first learner improved only from 34 to 31 days against a 14-day target. Accounts activating within 14 days churned at one-third the rate of other accounts. That is a promising intervention hypothesis, not proof: easier-to-activate customers may also be healthier for other reasons.

Growth generated activity, not expansion. Thirteen experiments and upgrade prompts accompanied a decline in seats added from 9,200 to 8,100. We cannot infer that those interventions caused the decline, but we have no basis for declaring them successful.

Some improvements are real; others require repair. Uptime reached 99.93%, a meaningful reliability improvement, though its retention effect is unmeasured. Completion rose from 41% to 58%, but the definition changed from finishing every module to finishing 80%. We must restate both periods under one definition before claiming improvement. Content volume and translations show capacity delivered, not that industry-fit problems were solved.

Proposed Q4 team OKRs

All baselines below are Q3 actuals. Delivery milestones belong in team plans, not as substitutes for these KRs.

Team and objectiveKey resultBaseline → Q4 target
Learner app: Make learning a sustained habit that earns renewal.Increase weekly active learners per paid seat, using the existing definition and consistent weekly aggregation.29% → 40%
Reduce quarterly logo churn, jointly with Admin; Learner leads adoption recovery in existing accounts.3.4% → 2.5%
Admin: Get customers to learning value quickly enough to stay.Reduce days from contract signature to first learner starting a course, retaining Q3’s aggregation method.31 → 14 days
Reduce quarterly logo churn, jointly with Learner; Admin leads rollout and administrator-value recovery.3.4% → 2.5%
Growth: Turn demonstrated customer value into durable expansion.Increase existing-account NRR; Growth owns the financial reconciliation, with retention contributions from Learner and Admin.102% → 108%
Increase seats added in existing accounts, retaining Q3’s counting definition.8,100 → 14,000

The churn KR is one shared result, not two independent contributions. Likewise, Growth’s NRR accountability does not make retention solely Growth’s responsibility. The CPO arbitrates cross-team trade-offs.

The behavioral targets recover ambitions already set in Q3; shipping the previous roadmap did not validate their feasibility. Teams should present intervention evidence and capacity assumptions before commitment. Learner and Admin should prioritize existing accounts with low adoption, stalled rollouts and upcoming renewals—not rely only on newly signed accounts whose retention impact may arrive later.

How these connect to the company goal

The causal hypothesis is:

Faster activation → sustained learning → stronger renewal value → fewer departures and more paid expansion.

The financial test is separate:

NRR = (opening recurring revenue + expansion − contraction − churned revenue) / opening recurring revenue.

Using the company’s existing measurement window, 108% means ending with $108 from each $100 of opening recurring revenue, versus Q3’s $102. Therefore, expansion less contraction and churn must improve by six percentage points of opening revenue.

We cannot honestly allocate that six-point improvement from the supplied data. Logo churn counts customers, not dollars; seats added omit pricing, discounts and offsetting seat reductions. A 0.9-point logo-churn improvement plus 5,900 additional seats does not automatically equal six NRR points.

In week one, Growth and Finance must build the account-level revenue bridge: opening revenue, renewal exposure, churn, contraction and expansion. That bridge must quantify a credible route to 108%; otherwise, revise the intervention plan rather than claim the seat target is sufficient.

Operating discipline

Review adoption and activation weekly; review renewal risk and the revenue bridge monthly. Freeze metric definitions, restate completion, and segment results by account size, tenure and renewal timing. Check that activation improvements are not just token first starts and that adoption gains are not driven by shrinking paid-seat denominators.

At quarter-end, report the two company outcomes first. No blended delivery score should again turn deteriorating retention into “a great quarter.”

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly63%
  • passAddresses the actual decision87%
  • passRespects explicit constraints22%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable60%
  • passReads the scorecard for what it is98%
  • passCatches the redefined key result99%
  • passKey results are outcomes, not output92%
  • passFocuses on the big rock64%
  • passShows how the goals add up56%
Run
Run
#1
Time to output
70 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

All got right 11

Uses the supplied evidence correctlyRightRight
Sonnet 5.5 · API

All claims about the current situation are directly from the supplied context or follow by arithmetic.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are directly supported by the supplied context or simple arithmetic.

Addresses the actual decisionRightRight
Sonnet 5.5 · API

Commits to clear Q4 OKRs with baselines and targets, and says what would change the plan (confirming company goal, data availability).

GPT-6 Astra · ChatGPT

The memo commits to a clear set of Q4 OKRs with baselines and targets, and specifies that a revenue bridge must show a credible route to 108% NRR or the plan should be revised.

Respects explicit constraintsRightRight
Sonnet 5.5 · API

Memo is under 1,100 words, addressed to Grace Mensah from Staff PM, and includes the requested sections.

GPT-6 Astra · ChatGPT

The output is a memo to Grace Mensah, within the 1,100-word limit, and includes the requested diagnosis and team OKRs with baselines and targets.

Identifies material uncertaintyRightRight
Sonnet 5.5 · API

Names assumptions (company goal, churn-to-NRR conversion, seat base) and asks for confirmations; provides a fallback if onboarding can't affect Q4 renewals.

GPT-6 Astra · ChatGPT

The output identifies key unknowns (causality of exit interviews, confounding factors in activation, inability to allocate NRR improvement from available data) and says how to resolve them with a revenue bridge.

Avoids unsupported claimsRightRight
Sonnet 5.5 · API

Interpretations like selection bias are labelled as possibilities ('could partly be'), and conclusions are drawn from the data.

GPT-6 Astra · ChatGPT

Interpretations and causal hypotheses are clearly labelled as such, and confident claims are limited to what the evidence establishes.

Produces the required deliverableRightRight
Sonnet 5.5 · API

Provides a complete memo with analysis of Q3, Q4 OKRs for the three teams with baselines and targets, and a ladder to the company goal.

GPT-6 Astra · ChatGPT

The memo is complete, in the right form for the CPO, and a PM could act on it with only light edits.

Reads the scorecard for what it isRightRight
Sonnet 5.5 · API

Splits output KRs (all 1.00) from outcome KRs (average 0.48) and contrasts with the missed company goals (NRR 102%, churn 3.4%).

GPT-6 Astra · ChatGPT

The memo splits scores into output (all 1.00) and outcome (~0.48 average) and contrasts them with the missed company goals, rather than reporting the 0.81 average.

Catches the redefined key resultRightRight
Sonnet 5.5 · API

Catches the redefined completion metric (baseline and actual not comparable) and the mislabelled 'Launch 3 HRIS integrations' as output.

GPT-6 Astra · ChatGPT

It catches both the redefined completion metric (not comparable) and the mislabeled HRIS integrations KR (output, not outcome), and explains the implications.

Key results are outcomes, not outputRightRight
Sonnet 5.5 · API

All Q4 key results are outcome metrics (weekly active learners, churn points, days to first learner, seats added) with baselines and targets; no shipping KRs.

GPT-6 Astra · ChatGPT

All proposed key results are measurable changes in customer behavior (adoption, activation, expansion, churn) with baselines and targets; shipping work is explicitly relegated to team plans.

Focuses on the big rockRightRight
Sonnet 5.5 · API

Focuses on usage and activation, drops content and budget as unaddressable this quarter, and gives each team a single objective with few KRs.

GPT-6 Astra · ChatGPT

Each team has one objective and at most two key results, focusing on the critical levers for retention and expansion, and output KRs are dropped.

Shows how the goals add upRightRight
Sonnet 5.5 · API

Links each KR to churn or NRR: non-use churn shared between Learner and Admin, HRIS churn to Admin, seats added to NRR, with arithmetic showing how they add up to the company goal.

GPT-6 Astra · ChatGPT

The memo provides a causal chain linking activation, adoption, and expansion to the company NRR and churn goals, and flags the shared churn KR and the need for a revenue bridge.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT98.5100.03None
2GPT-6.1 SolwithAPI98.5100.03None
3Sonnet 5.5withAPI89.494.43None
4GPT-6 LunawithAPI90.986.13None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review