Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 20 graded outputs by 7 models. 65% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Pathwise, working for the CPO, Grace Mensah. Last quarter's OKR scorecard is below, along with the company goal it was meant to serve. Grace wants a memo of no more than 1,100 words: what last quarter's OKRs actually tell us (not just the scores), and next quarter's OKRs for the Learner app, Admin and Growth teams, each with baselines and targets, that add up to the company goal.\n\nThe files are attached.

What the model was given5 items: About Pathwise, company_okr_q3.md, okr_scorecard_q3.csv (team_label is the type each team gave its key result), metric_dictionary.md, churn_reasons_q3.md
About PathwiseOnline learning for companies: courses employees take, and tools for the HR and L&D teams who buy it. Sold per seat on annual contracts.
company_okr_q3.md7 lines · Download
# Company OKR, Q3 (set by the CEO)

Objective: grow by keeping and expanding the customers we have.
- Net revenue retention: from 103% to 108%. Actual: 102%.
- Quarterly logo churn: from 3.1% to 2.5%. Actual: 3.4%.

Q3 average key result score across product teams: 0.81. The all-hands slide said "a great quarter".
okr_scorecard_q3.csv (team_label is the type each team gave its key result)team,objective,key_result,team_label,baseline,target,actual,score Content,Best course library in L&D,Publish 40 new courses,output,0,40,41,1.00 Content,Best course library in L&D,Course completion rate,outcome,41,55,58,1.00 Content,Best course library in L&D,Translate top 20 courses into 3 languages,output,0,60,60,1.00 Learner app,Learning people love,Ship mobile offline mode,output,0,1,1,1.00 Learner app,Learning people love,Weekly active learners per seat,outcome,0.31,0.4,0.29,0.00 Learner app,Learning people love,Ship streaks and badges,output,0,1,1,1.00 Learner app,Learning people love,App store rating,outcome,4.2,4.6,4.5,0.75 Admin,Admins get value fast,Launch 3 HRIS integrations,outcome,0,3,3,1.00 Admin,Admins get value fast,Days from contract to first learner,outcome,34,14,31,0.15 Admin,Admins get value fast,Ship new reporting dashboard,output,0,1,1,1.00 Growth,Expand inside accounts,Seats added in existing accounts,outcome,9200,14000,8100,0.00 Growth,Expand inside accounts,Run 12 expansion experiments,output,0,12,13,1.00 Growth,Expand inside accounts,Launch in-app upgrade prompts,output,0,1,1,1.00 Platform,Reliable at scale,Uptime,outcome,99.5,99.9,99.93,1.00 Platform,Reliable at scale,Migrate to new video CDN,output,0,1,1,1.00 Platform,Reliable at scale,Close 150 bugs,output,0,150,162,1.00
metric_dictionary.md5 lines · Download
# Metric dictionary (excerpt)

- Course completion rate: share of course enrolments marked complete. Until 31 July "complete" meant every module finished. From 1 August it means 80% of modules finished (changed so completion matches the certificate rule).
- Weekly active learners per seat: learners who start or continue a course in a week, divided by paid seats.
- Days from contract to first learner: calendar days from contract signature to the first learner starting a course.
churn_reasons_q3.md8 lines · Download
# Why accounts left in Q3 (exit interviews, 38 accounts)

- "Our people didn't use it": 17
- Budget cut: 9
- Moved to the LMS that came with our HR system: 7
- Content didn't fit our industry: 5

…
What a strong answer doesThe answer key the graders mark against

Reads the scorecard for what it is: the average is 0.81, but every output key result scored 1.00 (things shipped) while outcome key results averaged about 0.48, and the company missed both goals (NRR 102% against 108%, churn 3.4% against 2.5%). Two scores flatter: 'Launch 3 HRIS integrations' is labelled an outcome but is output, and course completion only 'beat' its target because completion was redefined on 1 August (all modules to 80%), so the result can't be compared with its baseline. The outcome key results that matter for churn moved backwards or barely (weekly active learners per seat 0.31 to 0.29; days to first learner 34 to 31; seats added 8,100 against 9,200). Builds next quarter's OKRs on the churn evidence: low usage is the top exit reason, and accounts reaching a first learner within 14 days churn at a third of the rate. So Admin owns days from contract to first learner (from 31), Learner app owns weekly active learners per seat (from 0.29), Growth owns seats added in existing accounts or net expansion, each with a baseline from the files and a target, with shipping work as initiatives. Notes the completion metric needs restating on one definition before anyone uses it.

Critical failures (cap the score at 40)
  • Describes last quarter as a success on the strength of the 0.81 average
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, data files, staff-level

The outputs

6 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit2 of 11 checks passed

The memo invents baselines for new metrics not in the supplied data, and includes output key results that are things to ship, requiring rework of the OKRs and verification of baselines before use.

Output

MEMORANDUM

TO: Grace Mensah, Chief Product Officer FROM: Staff PM, Product DATE: October 2, 2023 SUBJECT: Q3 OKR Post-Mortem & Q4 Product OKRs

---

Part 1: What Q3 Actually Tells Us (Beyond the 0.81 Score)

On paper, Q3 looked like a triumph. The company-wide average key result score across product teams was 0.81, prompting celebratory remarks at the all-hands. However, a closer look at the data reveals a stark disconnect between our product output and the company’s core objective: growth by keeping and expanding the customers we have.

The company's Q3 goals failed. Net Revenue Retention (NRR) dropped from our baseline of 103% to an actual of 102% (target: 108%), and quarterly logo churn rose from 3.1% to 3.4% (target: 2.5%).

Why did we miss our financial and retention targets while hitting an 0.81 KR score? The data points to a systemic structural flaw in how we set and measure OKRs: we rewarded teams for shipping outputs, while outcomes that actually drive retention and expansion flatlined or missed.

1. The Output-Outcome Illusion: Look at the team scorecards. Every single output KR scored 1.00. Content published 41 courses (target 40), translated 60 items, shipped mobile offline mode, shipped streaks and badges, launched three HRIS integrations, shipped a new reporting dashboard, ran 13 expansion experiments, launched in-app upgrade prompts, migrated a video CDN, and closed 162 bugs. We were exceptionally busy. Conversely, our critical outcomes largely failed: * Learner App: Weekly active learners per seat plummeted from 0.31 to an actual of 0.29 (target: 0.40, score: 0.00). * Admin: Days from contract to first learner barely budged, moving from 34 to 31 days (target: 14 days, score: 0.15). * Growth: Seats added in existing accounts fell from a baseline of 9,200 to 8,100 (target: 14,000, score: 0.00).

2. The Root Causes of Churn: Our Q3 exit interviews across 38 churned accounts provide the smoking gun. The leading reason for churn—cited by nearly half (17 out of 38)—was simple: "Our people didn't use it." Furthermore, the data reveals a powerful operational lever we have ignored: accounts that reached their first learner within 14 days of contract signature churned at one-third of the rate of the rest. Yet, our Admin team’s initiative to reduce this time stalled at 31 days.

3. Metric Distortions: Even our apparent wins require scrutiny. Content’s course completion rate hit 1.00 (rising from 41% to 58%), but the metric dictionary reveals a rule change on August 1st: the definition of "complete" was quietly lowered from 100% of modules finished to 80%. The "success" is largely an artifact of moving the goalposts.

The Takeaway: In Q3, product engineering functioned as an effective feature factory, but those features did not translate into user engagement, rapid time-to-value, or account expansion. To fix NRR and logo churn in Q4, we must completely realign product incentives away from output delivery and lock them onto core engagement and adoption metrics.

---

Part 2: Q4 OKRs for Learner App, Admin, and Growth Teams

To directly support the CEO's ongoing company objective—grow by keeping and expanding the customers we have—we are restructuring Q4 OKRs. Every team’s objectives are anchored to lagging retention drivers (engagement, time-to-value, and expansion).

#### 1. Learner App Team * Context & Rationale: With 17 out of 38 churned accounts citing lack of employee usage, the Learner App team must pivot away from vanity features (streaks and badges shipped in Q3 did nothing to stop the drop) and focus entirely on driving habitual weekly usage. * Objective: Drive habitual engagement to secure employee adoption and prevent account-level churn. * Key Results: * KR 1 (Outcome): Increase weekly active learners per seat from 0.29 to 0.36. (Baseline: 0.29, Target: 0.36) * KR 2 (Outcome): Increase 30-day course retention rate for newly enrolled learners from 22% to 35%. (Baseline: 22%, Target: 35%) * KR 3 (Output): Ship personalized "Next Step" AI recommendations on the home screen to streamline the path from login to learning. (Baseline: 0, Target: 1)

#### 2. Admin Team * Context & Rationale: Exit data proves that getting the first learner active within 14 days slashes churn by two-thirds. In Q3, we focused on HRIS integrations (output), but customers still languished for 31 days before their employees started learning. Admin must own the friction reduction between contract signature and organizational activation. * Objective: Accelerate time-to-value so customer organizations realize immediate ROI. * Key Results: * KR 1 (Outcome): Reduce days from contract signature to first learner from 31 days to 14 days. (Baseline: 31, Target: 14) * KR 2 (Outcome): Increase self-serve HRIS integration completion rate (admins successfully connecting without support intervention) from 40% to 75%. (Baseline: 40%, Target: 75%) * KR 3 (Output): Launch automated "Nudge Admin" workflow triggers when zero learners are active 7 days post-contract. (Baseline: 0, Target: 1)

#### 3. Growth Team * Context & Rationale: Q3 saw expansion seats drop to 8,100 against a 14,000 target. Running 12 experiments and launching upgrade prompts generated output, but no revenue expansion. Growth must transition from running experiments to driving net-new seat expansion inside existing healthy accounts. * Objective: Expand footprint and seat count within existing healthy accounts. * Key Results: * KR 1 (Outcome): Increase net new seats added in existing accounts from 8,100 (Q3 actual) to 12,500. (Baseline: 8,100, Target: 12,500) * KR 2 (Outcome): Increase conversion rate of in-app upgrade prompts from trial/capped tiers to paid seat expansions from 1.2% to 3.5%. (Baseline: 1.2%, Target: 3.5%) * KR 3 (Output): Launch automated team-invitation workflows for departmental managers in accounts nearing seat capacity limits. (Baseline: 0, Target: 1)

Claims the judge couldn’t find in the brief

  • 30-day course retention rate for newly enrolled learners baseline is 22%
  • Self-serve HRIS integration completion rate baseline is 40%
  • Conversion rate of in-app upgrade prompts baseline is 1.2%
  • Streaks and badges shipped in Q3 did nothing to stop the drop in weekly active learners
  • Product engineering functioned as an effective feature factory

Grades and run details

Decision model 41 · LLM judge 3 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly86%
  • passAddresses the actual decision42%
  • failRespects explicit constraints12%
  • failIdentifies material uncertainty36%
  • failAvoids unsupported claims36%
  • partialProduces the required deliverable48%
  • passReads the scorecard for what it is98%
  • partialCatches the redefined key result95%
  • partialKey results are outcomes, not output55%
  • partialFocuses on the big rock30%
  • partialShows how the goals add up87%
Run
Run
#1
Time to output
8 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is evidence-based, correctly reads the scorecard, and provides actionable OKRs; a PM could use it with only a quick edit to fill in the revenue baselines once Finance provides them.

Output

To: Grace Mensah, CPO From: Staff PM Subject: Q3 delivered features, not retained growth; proposed Q4 OKRs

What Q3 actually tells us

We missed the company objective despite a 0.81 product score. NRR fell from 103% to 102%, versus a 108% target. Quarterly logo churn rose from 3.1% to 3.4%, versus a 2.5% target. Calling this “a great quarter” confuses delivery with business impact.

The scorecard explains that disconnect:

  • Shipping dominated scoring. Ten of the 16 KRs measured outputs, including “Launch 3 HRIS integrations,” incorrectly labeled an outcome. All ten scored 1.00. The six genuine outcome KRs averaged approximately 0.48, not 0.81.
  • Learner value weakened while features shipped. Weekly active learners per paid seat fell from 0.31 to 0.29 despite offline mode, streaks and badges. A better app-store rating does not demonstrate adoption across employer accounts.
  • Admin setup barely improved. Contract-to-first-learner time moved from 34 to 31 days, far short of 14. Shipping integrations and reporting did not establish that customers reached value faster.
  • Expansion weakened despite experimentation. Seats added in existing accounts fell from 9,200 to 8,100—about 12%—while experiments and upgrade prompts exceeded delivery targets. Seats are also not revenue: discounts, seat removals and account losses can offset additions.
  • The content “win” is not comparable. Completion rose from 41% to 58%, but the definition changed from finishing every module to finishing 80%. We cannot separate real improvement from reclassification without recalculating both periods under one definition.
  • Reliability genuinely improved. Uptime exceeded target. That is valuable, but insufficient evidence of retention or expansion impact.

The churn evidence suggests where to intervene, without proving causality. Low usage was cited by 17 of 38 interviewed departing accounts, the largest category. Accounts reaching a first learner within 14 days churned at one-third the rate of other accounts. These findings support prioritizing activation and adoption; they do not establish that accelerating activation will itself cut churn by two-thirds. Customer characteristics may explain part of the association.

Q4 company commitment

Retain the CEO’s destination, rebased to Q3 actuals:

Company KRBaselineQ4 target
------:---:
Net revenue retention102%108%
Quarterly logo churn3.4%2.5%

Before scoring starts, Finance must confirm consistent NRR periods, cohort rules and revenue treatment. Annual contracts make renewal timing especially important.

Proposed team OKRs

Feature launches become initiatives, not scored KRs. Each team owns customer or commercial results.

Team / objectiveKey resultBaselineQ4 target
---------:---:
Learner app: Make paid seats deliver sustained learning valueWeekly active learners per paid seat, using the existing definition0.290.40
Admin: Get customers to value quickly and protect renewal revenueDays from contract to first learner, using the same aggregation as Q331 days14 days
AdminExisting-customer revenue lost to churn and contraction, divided by opening cohort revenueL₃, to be calculated from Q3 billingL₃ − 3 percentage points
Admin, shared with Learner and CSQuarterly logo churn3.4%2.5%
Growth: Turn demonstrated customer value into paid expansionExisting-customer expansion revenue, divided by opening cohort revenueE₃, to be calculated from Q3 billingE₃ + 3 percentage points
GrowthSeats added in existing accounts, retaining Q3’s definition8,10014,000

The logo-churn KR is one shared result, not three independently credited achievements. Admin coordinates the retention plan with CS; Learner owns adoption improvements. Growth’s revenue KR takes precedence over its seat-volume KR: buying seat growth through discounts is not success.

The revenue baselines are deliberately not invented. The supplied scorecard lacks expansion revenue, contraction revenue, churned revenue and opening recurring revenue. Finance must populate E₃ and L₃ before final approval. If Q3 revenue loss was below three percentage points, the proposed loss reduction is impossible; we must shift that portion of the six-point improvement budget to expansion and reassess feasibility.

How these add up

Using a consistent existing-customer cohort definition:

NRR = 100% + expansion revenue rate − churn/contraction revenue rate.

Q3 therefore implies E₃ − L₃ = 2 percentage points. The proposed Q4 improvement budget is:

  • Growth: three additional percentage points of expansion;
  • Admin, supported by Learner and CS: three fewer percentage points of revenue loss.

Together: 100% + (E₃ + 3) − (L₃ − 3) = 108%.

This is a six-point commercial bridge, not a claim that engagement or activation mechanically converts into revenue. Logo churn remains a separate constraint: retaining small accounts could improve logo churn without materially improving NRR. Conversely, one large expansion could improve NRR while many customers leave.

Operating discipline

In week one, Finance, Product and CS should establish the revenue baselines, renewal exposure and account-level retention/expansion plan. Validate that the improvement budget is achievable before locking the OKRs.

Review weekly by account cohort, not just company averages. Track activation coverage alongside time-to-first-learner so one early learner cannot conceal a stalled rollout. Inspect engagement across accounts so a few large customers cannot conceal widespread non-use.

Recalculate completion under a constant definition. Keep uptime as a guardrail. Do not average feature delivery and customer outcomes into another celebratory score: Q4 is successful only if customer value improves and the company retention goals are met.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly62%
  • passAddresses the actual decision90%
  • passRespects explicit constraints20%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims78%
  • partialProduces the required deliverable72%
  • passReads the scorecard for what it is100%
  • passCatches the redefined key result100%
  • passKey results are outcomes, not output74%
  • passFocuses on the big rock58%
  • passShows how the goals add up66%
Run
Run
#1
API response time
62 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output invents baselines for new metrics (30-day course retention 22%, self-serve HRIS integration completion 40%, in-app upgrade prompt conversion 1.2%) that are not in the supplied context.

GPT-6.1 Sol · API

All factual claims about the current situation are taken directly from the supplied context or derived by arithmetic; no invented numbers or facts.

Addresses the actual decisionMixedRight
Gemini 3.5 Flash-Lite · Gemini

The memo does not state what result or condition would change the recommended OKRs, failing to meet the requirement to say what would change the answer.

GPT-6.1 Sol · API

The memo commits to a clear set of Q4 OKRs with baselines and targets, addresses the CPO, and specifies conditions (revenue baselines, feasibility) that would change the plan.

Respects explicit constraintsMixedRight
Gemini 3.5 Flash-Lite · Gemini

The output is a memo under 1,100 words, addressed to Grace Mensah, and includes the requested analysis and next-quarter OKRs with baselines and targets.

GPT-6.1 Sol · API

The output is a memo to Grace Mensah, well under 1,100 words, and respects the requested form and reader.

Identifies material uncertaintyWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output names no unknowns that could change the decision and does not say how they would be resolved.

GPT-6.1 Sol · API

It names missing revenue baselines, the redefined completion metric, and the possibility that Q3 loss is too low to hit the target, and says how to resolve each.

Avoids unsupported claimsWrongRight
Gemini 3.5 Flash-Lite · Gemini

Causal claims like 'streaks and badges did nothing to stop the drop' and 'product engineering functioned as an effective feature factory' are presented as fact without being labelled as hypotheses.

GPT-6.1 Sol · API

Interpretations and causal claims are clearly labelled as suggestions or hypotheses, not presented as established fact.

Catches the redefined key resultWrongRight
Gemini 3.5 Flash-Lite · Gemini

The memo catches the course completion redefinition but does not mention that the 'Launch 3 HRIS integrations' KR is output mislabelled as outcome.

GPT-6.1 Sol · API

It catches both the redefined completion metric (definition change on 1 August) and the integrations KR being output labelled as outcome.

Key results are outcomes, not outputWrongRight
Gemini 3.5 Flash-Lite · Gemini

Each team's third key result is a thing to ship (e.g., 'Ship personalized AI recommendations'), not a measurable change in customer or business behaviour.

GPT-6.1 Sol · API

Every proposed key result is a measurable change in customer or business behaviour with a baseline and target; shipping work is explicitly moved to initiatives.

Focuses on the big rockWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output does not state what was dropped from the previous quarter's scope or why, failing the requirement to cut goals down to the few that matter most.

GPT-6.1 Sol · API

The OKRs are cut to a few critical outcomes per team, dropping output KRs and focusing on activation, time-to-value, expansion revenue, and churn.

Shows how the goals add upWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output key results (shipping features) are not linked to a company goal with reasoning or data, and no KR is flagged as not serving a company goal.

GPT-6.1 Sol · API

The memo shows how each team KR contributes to NRR and logo churn via a revenue bridge formula, and flags the separate logo churn constraint.

All got right 2

Produces the required deliverableRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo is complete, in the right form, within the word limit, and a PM could act on its structure with light edits.

GPT-6.1 Sol · API

The memo provides the required analysis of Q3 and proposed Q4 OKRs for the three teams, with baselines and targets, and is usable as-is.

Reads the scorecard for what it isRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo explicitly splits output KRs (all scored 1.00) from outcome KRs (most missed) and contrasts them with the missed company goals.

GPT-6.1 Sol · API

It splits the scores into output (all 1.00) and outcome (average ~0.48) and contrasts them with the missed company goals.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT98.5100.03None
2GPT-6.1 SolwithAPI98.5100.03None
3Sonnet 5.5withAPI89.494.43None
4GPT-6 LunawithAPI90.986.13None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review