Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 20 graded outputs by 7 models. 65% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Pathwise, working for the CPO, Grace Mensah. Last quarter's OKR scorecard is below, along with the company goal it was meant to serve. Grace wants a memo of no more than 1,100 words: what last quarter's OKRs actually tell us (not just the scores), and next quarter's OKRs for the Learner app, Admin and Growth teams, each with baselines and targets, that add up to the company goal.\n\nThe files are attached.

What the model was given5 items: About Pathwise, company_okr_q3.md, okr_scorecard_q3.csv (team_label is the type each team gave its key result), metric_dictionary.md, churn_reasons_q3.md
About PathwiseOnline learning for companies: courses employees take, and tools for the HR and L&D teams who buy it. Sold per seat on annual contracts.
company_okr_q3.md7 lines · Download
# Company OKR, Q3 (set by the CEO)

Objective: grow by keeping and expanding the customers we have.
- Net revenue retention: from 103% to 108%. Actual: 102%.
- Quarterly logo churn: from 3.1% to 2.5%. Actual: 3.4%.

Q3 average key result score across product teams: 0.81. The all-hands slide said "a great quarter".
okr_scorecard_q3.csv (team_label is the type each team gave its key result)team,objective,key_result,team_label,baseline,target,actual,score Content,Best course library in L&D,Publish 40 new courses,output,0,40,41,1.00 Content,Best course library in L&D,Course completion rate,outcome,41,55,58,1.00 Content,Best course library in L&D,Translate top 20 courses into 3 languages,output,0,60,60,1.00 Learner app,Learning people love,Ship mobile offline mode,output,0,1,1,1.00 Learner app,Learning people love,Weekly active learners per seat,outcome,0.31,0.4,0.29,0.00 Learner app,Learning people love,Ship streaks and badges,output,0,1,1,1.00 Learner app,Learning people love,App store rating,outcome,4.2,4.6,4.5,0.75 Admin,Admins get value fast,Launch 3 HRIS integrations,outcome,0,3,3,1.00 Admin,Admins get value fast,Days from contract to first learner,outcome,34,14,31,0.15 Admin,Admins get value fast,Ship new reporting dashboard,output,0,1,1,1.00 Growth,Expand inside accounts,Seats added in existing accounts,outcome,9200,14000,8100,0.00 Growth,Expand inside accounts,Run 12 expansion experiments,output,0,12,13,1.00 Growth,Expand inside accounts,Launch in-app upgrade prompts,output,0,1,1,1.00 Platform,Reliable at scale,Uptime,outcome,99.5,99.9,99.93,1.00 Platform,Reliable at scale,Migrate to new video CDN,output,0,1,1,1.00 Platform,Reliable at scale,Close 150 bugs,output,0,150,162,1.00
metric_dictionary.md5 lines · Download
# Metric dictionary (excerpt)

- Course completion rate: share of course enrolments marked complete. Until 31 July "complete" meant every module finished. From 1 August it means 80% of modules finished (changed so completion matches the certificate rule).
- Weekly active learners per seat: learners who start or continue a course in a week, divided by paid seats.
- Days from contract to first learner: calendar days from contract signature to the first learner starting a course.
churn_reasons_q3.md8 lines · Download
# Why accounts left in Q3 (exit interviews, 38 accounts)

- "Our people didn't use it": 17
- Budget cut: 9
- Moved to the LMS that came with our HR system: 7
- Content didn't fit our industry: 5

…
What a strong answer doesThe answer key the graders mark against

Reads the scorecard for what it is: the average is 0.81, but every output key result scored 1.00 (things shipped) while outcome key results averaged about 0.48, and the company missed both goals (NRR 102% against 108%, churn 3.4% against 2.5%). Two scores flatter: 'Launch 3 HRIS integrations' is labelled an outcome but is output, and course completion only 'beat' its target because completion was redefined on 1 August (all modules to 80%), so the result can't be compared with its baseline. The outcome key results that matter for churn moved backwards or barely (weekly active learners per seat 0.31 to 0.29; days to first learner 34 to 31; seats added 8,100 against 9,200). Builds next quarter's OKRs on the churn evidence: low usage is the top exit reason, and accounts reaching a first learner within 14 days churn at a third of the rate. So Admin owns days from contract to first learner (from 31), Learner app owns weekly active learners per seat (from 0.29), Growth owns seats added in existing accounts or net expansion, each with a baseline from the files and a target, with shipping work as initiatives. Notes the completion metric needs restating on one definition before anyone uses it.

Critical failures (cap the score at 40)
  • Describes last quarter as a success on the strength of the 0.81 average
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, data files, staff-level

The outputs

6 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 11 checks passed

The memo is clear, evidence-based, and actionable; a PM could use it with minor wording tweaks to emphasize the output/outcome score split.

Output

Memo

To: Grace Mensah, CPO Subject: What Q3 OKRs tell us—and Q4 OKRs tied to retention and expansion

What Q3 tells us

Q3 was not a “great quarter” against the company goal. Net revenue retention fell to 102%, below the 108% goal, and logo churn rose to 3.4%, above the 2.5% goal. The 0.81 average team score obscures that miss: many high-scoring KRs measured shipping or activity, not customer value.

The clearest warning is usage. Learner app weekly active learners per paid seat declined from 0.31 to 0.29, missing its 0.40 target. In exit interviews, “our people didn’t use it” was the most common reason for leaving (17 of 38 accounts). That makes engagement a credible retention lever to investigate—not proof that low app usage caused those departures.

Onboarding is another promising lever: time to first learner was 31 days, versus a 14-day target. Accounts reaching a first learner within 14 days churned at one-third the rate of the rest. That association supports prioritizing faster activation, but does not establish causation. Seven of the 38 departing accounts moved to an LMS bundled with their HR system, so integrations may also matter; Q3’s three launches are an output, not evidence of customer adoption or retention impact.

Expansion also underperformed: existing-account seats added were 8,100, below both the 14,000 target and the 9,200 baseline. Running 13 experiments and shipping upgrade prompts did not translate into the intended result. Meanwhile, Content’s completion result is hard to interpret: the definition changed on 1 August, from every module to 80% of modules, so 58% cannot be fairly compared with the 55% target without a consistent-definition view. Strong uptime and course delivery are useful foundations, but do not offset the missed retention and expansion outcomes.

Proposed Q4 OKRs

Keep the company outcome explicit and shared: NRR 102% → 108%; quarterly logo churn 3.4% → 2.5%. These are not numbers that team KRs can literally be summed to produce. They are the outcomes the teams collectively support; the team KRs below are leading indicators. Review company outcomes alongside them, using consistent cohort and revenue definitions.

Team / objectiveKey result: Q3 baseline → Q4 target
Learner app — Make learning a regular habit for more paid learnersWeekly active learners per paid seat: 0.29 → 0.40
Admin — Get new customers to learner value soonerCalendar days from contract signature to first learner starting a course: 31 → 14 or fewer
Growth — Expand paid seats in existing accountsPaid seats added in existing accounts: 8,100 → 14,000

Use the existing metric definitions and reporting windows for these comparisons; do not count shipped features, experiments, or prompts as substitutes for the outcomes. The Q3 data provided does not show how expansion seats translate into NRR, or how team-level changes affect logo churn. For Q4 reviews, Finance and the teams should make the NRR bridge visible—expansion, contraction, and churn—and check whether faster activation and higher learner activity precede improved retention. That will let us judge whether these team results are actually adding up to the company goal.

Grades and run details

Decision model 86 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly83%
  • passAddresses the actual decision50%
  • passRespects explicit constraints52%
  • partialIdentifies material uncertainty26%
  • passAvoids unsupported claims87%
  • passProduces the required deliverable45%
  • partialReads the scorecard for what it is85%
  • partialCatches the redefined key result32%
  • passKey results are outcomes, not output99%
  • passFocuses on the big rock75%
  • passShows how the goals add up32%
Run
Run
#1
API response time
24 s
Submitted
5 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 1

Reads the scorecard for what it isWrong

Does not split the scores by output and outcome with numbers; only says 'many high-scoring KRs measured shipping or activity' without quantifying the perfect output scores versus poor outcome scores.

Got right 10

Uses the supplied evidence correctlyRight

All factual claims about the current situation are directly from the supplied context or arithmetic.

Addresses the actual decisionRight

Commits to specific Q4 OKRs with baselines and targets, framed for the CPO, and suggests what to monitor to validate them.

Respects explicit constraintsRight

Memo format, under 1,100 words, addresses the three requested teams and ties to the company goal.

Identifies material uncertaintyRight

Identifies causation uncertainty, metric redefinition, and missing NRR bridge, and says how to resolve them in Q4 reviews.

Avoids unsupported claimsRight

Hypotheses are clearly labelled as not proven, and interpretations are cautious.

Produces the required deliverableRight

Provides a complete memo with analysis and next-quarter OKRs as requested, usable as is.

Catches the redefined key resultRight

Catches both the redefined completion metric and the output-labelled-as-outcome integrations KR, and explains the implications for interpreting scores.

Key results are outcomes, not outputRight

All proposed key results are outcome metrics with baselines and targets; no shipping work appears as a KR.

Focuses on the big rockRight

Focuses on three teams with one KR each, implicitly dropping Content and Platform as not directly driving the retention goal.

Shows how the goals add upRight

Links each team KR to the company retention/expansion goal, acknowledges data gaps, and suggests how to verify the connection.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT98.5100.03None
2GPT-6.1 SolwithAPI98.5100.03None
3Sonnet 5.5withAPI89.494.43None
4GPT-6 LunawithAPI90.986.13None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review