Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Focuses on the big rock100% pass
    The output reduces the draft to one objective with three KRs, explicitly drops dark mode, AI summaries, and NPS, and explains why.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  3. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Fieldnote's onboarding squad. The squad drafted its OKRs for next quarter and asked you to tidy them up before they go to leadership on Friday. Rewrite them, and add a short note to the squad on what you changed and why. No more than 600 words in all. What we know is below.

What the model was given6 items: About Fieldnote, The company's goal for next quarter (set by leadership), The squad's draft, Data, A note from HR, The squad
About FieldnoteShared notes and docs for teams, sold self-serve. Teams sign up free and pay once they pass five members.
The company's goal for next quarter (set by leadership)Objective: make Fieldnote the place new teams get started. Key result: raise the share of new teams still active in week 4 from 34% to 40%.
The squad's draftObjective 1: Ship the new onboarding. Key results: launch an onboarding checklist; ship 5 templates; redesign the invite flow; launch onboarding emails. Objective 2: Improve quality. Key results: fix 50 bugs; get P1 bugs to zero. Objective 3: Delight users. Key results: raise NPS from 31 to 40; launch dark mode; ship AI summaries; run 10 user interviews; publish 4 help articles.
DataNew teams where three or more people join in the first week: 58% still active in week 4. Everyone else: 21%. Today 29% of new teams get three or more people in the first week. Median time from sign-up to the first note someone else opens: 3 days.
A note from HRFrom this year, each squad member's bonus multiplier is set by their squad's OKR score.
The squadFive engineers, one designer and you.
What a strong answer doesThe answer key the graders mark against

Rewrites the draft into one objective tied to the company's (new teams get started) with two or three outcome key results built from the data, such as raising the share of new teams with three or more people in week one from 29% to a stated target, and cutting the time to the first shared note from 3 days to a stated target, with the company's week-4 retention as the result they serve. The checklist, templates, invite redesign and emails become initiatives, not key results. Drops what doesn't serve the goal (dark mode, AI summaries) or moves it out of the OKRs; keeps bug work as a health measure rather than an objective. Flags that tying bonuses to OKR scores will push the squad to set safe targets, and suggests keeping them apart. Explains the changes to the squad briefly and warmly.

Critical failures (cap the score at 40)
  • Keeps the draft's key results essentially unchanged
Task

Mid-level · v1.0 · synthetic · B2B SaaS, self-serve, goals

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 11 checks passed

The output is clear and actionable; the unsupported correlation claim can be fixed with a quick wording edit, and the rest is solid.

Output

Onboarding squad OKRs, next quarter

Objective 1: New teams get their teammates in during week 1, and stay (supports the company goal: make Fieldnote the place new teams get started)

  • KR1: Raise the share of new teams with 3+ people in week 1 from 29% to 45%.
  • KR2: Cut the median time from sign-up to the first note someone else opens from 3 days to 1 day.
  • KR3: Raise the share of new teams active in week 4 from 34% to 40% (shared with the company KR).

Planned bets, not KRs: onboarding checklist, invite flow redesign, 5 templates, onboarding emails. We'll drop or swap any of them if they aren't moving KR1 and KR2.

Objective 2: Keep the product dependable while we change onboarding

  • KR1: Zero open P1 bugs at quarter end, and no P1 open longer than 5 days at any point.

Backlog, not this quarter: dark mode, AI summaries, NPS target, help articles.

---

Note to the squad

Thanks for the draft. Here's what I changed and why.

1. Outcomes, not deliverables. Most of the draft listed things we'd ship. We could ship all of it and leave week-4 activity flat. The checklist, invite redesign, templates and emails are now the bets we make to hit the KRs.

2. One clear line to the company goal. Teams where 3+ people join in week 1 are still active in week 4 58% of the time, versus 21% for everyone else, and only 29% of teams reach 3+ today. So KR1 is about getting more teams to that point. Reaching 40% overall probably needs about 45% of teams there, assuming the relationship holds. KR2 measures how quickly a new team sees value from a teammate.

3. Cut and moved. Dark mode, AI summaries and the NPS target have no clear link to the company goal, and NPS isn't something onboarding alone can move. They go to the backlog, not the bin. User interviews and help articles are useful, but as activities inside Objective 1, not KRs. "Fix 50 bugs" rewarded volume over impact, so it became a P1 guardrail.

4. Two caveats to settle before Friday. - The data implies a week-4 rate of about 32% (0.29×58% + 0.71×21%), not the 34% in the company KR. Likely a definition or timing difference, but I'll confirm which baseline we're using. - The 58% vs 21% gap is correlation. Teams that invite people quickly may already be more committed, so the real lift from nudging invites will probably be smaller. That's why KR3 sits next to KR1, so we'll see whether the outcome really moves.

5. The bonus link. Since scores now affect everyone's multiplier, the safe play is to pick easy, output-style KRs. I've deliberately not done that. KR1 and KR2 are stretch targets, and I'd rather be upfront about it than quietly sandbag. Before Friday I'll ask leadership and HR how scores convert to multipliers, and I'll push for 0.7 on stretch KRs to count as a good result. I'll share the answer as soon as I have it. If you think a target is unrealistic, tell me before Friday so we can adjust it openly.

Please reply with objections or number challenges by Thursday noon.

Claims the judge couldn’t find in the brief

  • The 58% vs 21% gap is correlation.
  • The safe play is to pick easy, output-style KRs.

Grades and run details

Decision model 95 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly59%
  • passAddresses the actual decision81%
  • partialRespects explicit constraints23%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable80%
  • passBuilds the key results on the data100%
  • passFlags the bonus link100%
  • passKey results are outcomes, not output73%
  • passFocuses on the big rock100%
  • passShows how the goals add up59%
Run
Run
#1
API response time
25 s
Submitted
5 Oct 2026

Needs more than a quick edit8 of 11 checks passed

The output fails to address the bonus-link risk properly and makes an unsupported claim about the biggest retention lever, requiring edits before use.

Output

To the Onboarding Squad:

Thanks for putting these together. I’ve rewritten our OKRs for leadership. Before I explain why, please note HR’s rule: our bonuses are now tied directly to our OKR scores. That means our OKRs must be about outcomes, not output, and they must tightly align with the company goal.

Here is what I changed:

  1. Dropped the "feature laundry list." Draft 1 listed outputs (checklists, dark mode, 50 bugs). If we ship all those and week 4 retention doesn't budge, leadership fails us—and our bonuses drop. We need to focus on what happens to the user, not what we build.
  2. Anchored on the Company Goal: Leadership wants week 4 team retention up from 34% to 40%. Our entire quarter must serve this.
  3. Used our Data: Our data shows a massive retention cliff: teams where 3+ people join in week one retain at 58%, versus 21% for everyone else. Right now, only 29% hit that 3+ threshold. Furthermore, the median time to a shared note is 3 days—too slow.

Here is the finalized version going to leadership on Friday:

*

Q3 Onboarding Squad OKRs

Objective: Get new teams collaborating faster so they stick around.

  • KR 1: Raise the share of new teams with 3+ active members in week 1 from 29% to 40%. (Directly attacks our biggest retention lever).
  • KR 2: Reduce the median time from sign-up to the first note opened by a second team member from 3 days to under 24 hours. (Accelerates "time-to-value" and teamwork).
  • KR 3: Raise the share of new teams still active in week 4 from 34% to 40%. (Our direct contribution to the company goal).

*

What this means for us: Features like the checklist, templates, AI summaries, and invite flow redesign are still our tools to hit these goals, but we’ll prioritize them based on whether they actually drive team invites and fast collaboration. Let’s crush this.

Grades and run details

Decision model 82 · LLM judge 8 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly76%
  • passAddresses the actual decision71%
  • passRespects explicit constraints56%
  • failIdentifies material uncertainty87%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable92%
  • passBuilds the key results on the data100%
  • partialFlags the bonus link93%
  • passKey results are outcomes, not output99%
  • passFocuses on the big rock100%
  • passShows how the goals add up79%
Run
Run
#1
Time to output
3 s
Submitted
5 Oct 2026

Needs more than a quick edit9 of 11 checks passed

The learning objective uses activity-based KRs instead of outcome metrics, and the bonus-link handling lacks a concrete fix; a PM would need to rework both.

Output

Onboarding squad OKRs: next quarter

Company goal we serve: Raise the share of new teams still active in week 4 from 34% to 40%.

Objective 1: Get new teams working together in their first week - KR1: Raise the share of new teams with 3+ people joining in week one from 29% to 45%. - KR2: Cut the median time from sign-up to the first note a teammate opens from 3 days to 1. - KR3 (shared with the company): Raise week-4 active new teams from 34% to 40%.

Objective 2: Learn what makes a new team stick - KR1: By week 4, interview 10 teams that stalled and 10 that stuck, and share the findings. - KR2: By mid-quarter, run a test showing whether adding invites causes better week-4 retention, or whether teams that invite were just more committed to begin with.

Guardrails (watched, not scored) - Zero open P1 bugs in sign-up, invite and first-note flows. - Free-to-paid conversion at the five-member mark doesn't drop.

Bets we'll try to move the KRs: the invite-flow redesign, the onboarding checklist, onboarding emails, and starter templates built for sharing. We'll drop any bet that doesn't move KR1 or KR2.

---

Note to the squad

Thanks for the draft. The work in it is good. Here's what I changed before Friday.

1. I swapped launches for results. "Ship the checklist" scores 100% even if nobody uses it. Leadership asked for more teams active in week 4, so we now measure the changes in behaviour that lead there. Your launches are still in the plan, as bets.

2. I built the plan around the strongest signal we have. Teams where 3+ people join in week one are 58% active at week 4. Everyone else is at 21%. Only 29% of teams reach 3+ today. Getting teammates in, and getting them to open each other's notes, is the main lever. That's why the invite redesign comes first.

3. I cut dark mode, AI summaries, help articles, the NPS target and "fix 50 bugs." Some of these are worth doing, but none clearly moves week-4 retention for new teams. Five engineers and one designer can't do them all and still do onboarding well. Quality is now a guardrail on the flows new teams actually touch.

4. I added a learning objective because our data may mislead us. Teams that invite people may stick because they were already serious, not because they invited. If so, pushing invites won't move retention much. A test early in the quarter tells us before we spend all of it.

5. Two things I'll raise with leadership: - The numbers don't add up. 29% × 58% + 71% × 21% gives about 32% week-4 retention, not 34%. The metrics may be defined differently. I'll check before Friday. - The gap is big. Even if invites cause retention, hitting 40% would mean about half of new teams reaching 3+ people, up from 29%. My 45% target for KR1 is already a stretch. I'll be upfront that the company KR probably needs other teams' help too.

6. The bonus link: I want to be straight with you. HR now sets each of our bonus multipliers from our OKR score. That rewards safe targets. The old draft was mostly launches we control, so it would likely have scored well. The new one has ambitious targets we may only partly hit. Good OKRs often land around 70%.

I don't think the answer is to lower the bar. Easy targets would make our OKRs useless to leadership and to us. I'll ask leadership and HR how a 70% score on a stretch target will be read for bonuses. I'll come back to you before we lock anything. If the answer is bad, we'll decide together how to handle it.

Happy to talk any of this through before Friday.

Grades and run details

Decision model 91 · LLM judge 9 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly74%
  • passAddresses the actual decision77%
  • partialRespects explicit constraints24%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims29%
  • passProduces the required deliverable87%
  • passBuilds the key results on the data100%
  • passFlags the bonus link98%
  • partialKey results are outcomes, not output68%
  • passFocuses on the big rock100%
  • passShows how the goals add up67%
Run
Run
#1
Time to output
27 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyMixedRightRight
Sonnet 5.5 · API

Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.

Gemini 3.5 Flash-Lite · Gemini

All factual claims about the current situation are directly from the brief or supplied context.

Opus 5.5 · Claude

All statements about the current situation are taken directly from the brief or derived by correct arithmetic.

Identifies material uncertaintyRightWrongRight
Sonnet 5.5 · API

Names baseline discrepancy and correlation vs causation, and says how to resolve them; also addresses bonus conversion uncertainty.

Gemini 3.5 Flash-Lite · Gemini

The output does not name any unknowns that could change the decision or how they would be resolved.

Opus 5.5 · Claude

It identifies the causation vs correlation risk with invites and the discrepancy in the numbers, and says how to resolve them.

Avoids unsupported claimsMixedWrongRight
Sonnet 5.5 · API

Presents 'the gap is correlation' and 'the safe play is to pick easy KRs' as facts without labeling them as hypotheses.

Gemini 3.5 Flash-Lite · Gemini

The claim 'biggest retention lever' is presented as fact without support from the data, which only shows correlation.

Opus 5.5 · Claude

Hypotheses like 'teams that invite may stick because they were already serious' are clearly labelled as possibilities.

Flags the bonus linkRightWrongMixed
Sonnet 5.5 · API

Flags risk of safe targets due to bonus link and suggests pushing for 0.7 on stretch KRs to count as good.

Gemini 3.5 Flash-Lite · Gemini

The output notes the bonus-OKR link but does not suggest separating them or another concrete fix for the risk of safe targets.

Opus 5.5 · Claude

It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.

Key results are outcomes, not outputRightRightWrong
Sonnet 5.5 · API

All KRs are outcome metrics with baselines and targets; shipping work is listed as initiatives.

Gemini 3.5 Flash-Lite · Gemini

All key results are measurable changes in customer behavior, not things to ship.

Opus 5.5 · Claude

Objective 2's key results are activities (interviews, running a test) without baselines and targets, not measurable outcome metrics.

All got right 6

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

Commits to rewritten OKRs and note, and says squad feedback by Thursday could change targets.

Gemini 3.5 Flash-Lite · Gemini

The output commits to a clear set of rewritten OKRs and a note, addressing the request.

Opus 5.5 · Claude

The output commits to a clear set of rewritten OKRs and a note, as requested, without deferring the decision.

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

Output is a rewritten OKRs and note under 600 words, addressed to the squad.

Gemini 3.5 Flash-Lite · Gemini

The output is a note to the squad with the OKRs, well under 600 words, respecting the form and length.

Opus 5.5 · Claude

The output is a rewrite of the OKRs plus a note to the squad, and the total word count is under 600.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

Provides complete OKRs and note within word limit, usable by the squad.

Gemini 3.5 Flash-Lite · Gemini

The note and OKRs are complete, within length, and usable by the squad with light edits.

Opus 5.5 · Claude

The deliverable is a complete OKR rewrite and squad note, within the word limit, and usable as is.

Builds the key results on the dataRightRightRight
Sonnet 5.5 · API

KR1 uses the 29% baseline and KR2 uses the 3-day median, both from the data.

Gemini 3.5 Flash-Lite · Gemini

KR1 and KR2 are built on the first-week joining data and time to shared note, with baselines and targets.

Opus 5.5 · Claude

KR1 uses the 3+ people joining data with baseline 29% and target 45%; KR2 uses the time-to-first-shared-note data with baseline 3 days and target 1 day.

Focuses on the big rockRightRightRight
Sonnet 5.5 · API

Cuts to one main objective with three KRs and a quality guardrail, drops unrelated items and explains why.

Gemini 3.5 Flash-Lite · Gemini

The output cuts the goals to one objective with three KRs and explains what was dropped.

Opus 5.5 · Claude

It cuts the goals to two objectives, drops several items, and explains what was dropped and why.

Shows how the goals add upRightRightRight
Sonnet 5.5 · API

Links KR1 and KR2 to company goal via the data, and flags items that don't serve it.

Gemini 3.5 Flash-Lite · Gemini

Each KR is linked to the company goal with reasoning or data, and no KR serves no company goal.

Opus 5.5 · Claude

Every key result is linked to the company goal with the data that connects them, and the learning objective is tied to a risk that could undermine the main lever.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini77.362.52None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review