Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Focuses on the big rock100% pass
    The output reduces the draft to one objective with three KRs, explicitly drops dark mode, AI summaries, and NPS, and explains why.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  3. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Fieldnote's onboarding squad. The squad drafted its OKRs for next quarter and asked you to tidy them up before they go to leadership on Friday. Rewrite them, and add a short note to the squad on what you changed and why. No more than 600 words in all. What we know is below.

What the model was given6 items: About Fieldnote, The company's goal for next quarter (set by leadership), The squad's draft, Data, A note from HR, The squad
About FieldnoteShared notes and docs for teams, sold self-serve. Teams sign up free and pay once they pass five members.
The company's goal for next quarter (set by leadership)Objective: make Fieldnote the place new teams get started. Key result: raise the share of new teams still active in week 4 from 34% to 40%.
The squad's draftObjective 1: Ship the new onboarding. Key results: launch an onboarding checklist; ship 5 templates; redesign the invite flow; launch onboarding emails. Objective 2: Improve quality. Key results: fix 50 bugs; get P1 bugs to zero. Objective 3: Delight users. Key results: raise NPS from 31 to 40; launch dark mode; ship AI summaries; run 10 user interviews; publish 4 help articles.
DataNew teams where three or more people join in the first week: 58% still active in week 4. Everyone else: 21%. Today 29% of new teams get three or more people in the first week. Median time from sign-up to the first note someone else opens: 3 days.
A note from HRFrom this year, each squad member's bonus multiplier is set by their squad's OKR score.
The squadFive engineers, one designer and you.
What a strong answer doesThe answer key the graders mark against

Rewrites the draft into one objective tied to the company's (new teams get started) with two or three outcome key results built from the data, such as raising the share of new teams with three or more people in week one from 29% to a stated target, and cutting the time to the first shared note from 3 days to a stated target, with the company's week-4 retention as the result they serve. The checklist, templates, invite redesign and emails become initiatives, not key results. Drops what doesn't serve the goal (dark mode, AI summaries) or moves it out of the OKRs; keeps bug work as a health measure rather than an objective. Flags that tying bonuses to OKR scores will push the squad to set safe targets, and suggests keeping them apart. Explains the changes to the squad briefly and warmly.

Critical failures (cap the score at 40)
  • Keeps the draft's key results essentially unchanged
Task

Mid-level · v1.0 · synthetic · B2B SaaS, self-serve, goals

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 11 checks passed

The note does not flag the risk of safe targets from the bonus-OKR tie, and the unsupported claim about shared value would need rephrasing; a PM would have to add that risk flag and adjust the claim.

Output

Objective: Help new teams reach shared value quickly and make Fieldnote their place to get started.

Key results 1. Increase the share of new teams still active in week 4 from 34% to 40% (shared with the company goal). 2. Increase the share of new teams with three or more people joining in their first week from 29% to 40%. 3. Reduce the median time from sign-up to the first note opened by someone else from 3 days to 1 day.

Candidate initiatives: Test an onboarding checklist, templates, invite-flow changes and onboarding emails. Use interviews and support feedback to guide the work; treat bug fixes as necessary reliability work, not OKRs.

Note to the squad

I replaced launch counts and broad activity goals with measurable outcomes tied to the company’s week-four retention goal. The data suggests teams that bring in three or more people early are more likely to stick around, and getting a note opened by someone else is a concrete sign of shared value. The feature ideas may help achieve these results, but shipping them isn’t success by itself.

I left out NPS, dark mode, AI summaries, bug counts, interviews and help articles as KRs because they don’t directly show whether onboarding is working. They can still be useful work or learning activities. Since OKR scores affect bonuses, let’s agree on cohort definitions, measurement and ownership before the quarter starts—and review progress against the same definitions throughout.

Grades and run details

Decision model 91 · LLM judge 8 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly88%
  • passAddresses the actual decision58%
  • passRespects explicit constraints60%
  • partialIdentifies material uncertainty77%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable78%
  • passBuilds the key results on the data100%
  • partialFlags the bonus link57%
  • passKey results are outcomes, not output100%
  • passFocuses on the big rock100%
  • passShows how the goals add up59%
Run
Run
#1
API response time
16 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 11 checks passed

The output is clear, evidence-based, and directly usable by a PM with at most minor wording tweaks; it flags a data discrepancy and bonus risk proactively.

Output

Proposed onboarding OKRs — next quarter

Objective: Help new teams experience shared value quickly and keep collaborating.

Key resultBaseline → target
1. Increase the share of new teams still active in week 4. This is our contribution to the company’s retention goal, shared with other squads that influence it.34% → 40%
2. Increase the share of new teams with three or more people joining in their first week.29% → 40%
3. Reduce median time from sign-up to the first note opened by another team member. The share of teams reaching this milestone within seven days must not decline.3 days → 1 day

The targets for KRs 2 and 3 are proposed, pending a feasibility review before leadership approval. They are not forecasts of retention impact.

Measurement and quality guardrails

  • Freeze cohort eligibility, “active,” membership, and shared-note definitions before the quarter starts. Use consistent definitions for baselines and results.
  • Score week-4 retention only on cohorts that have reached week 4; report immature cohorts separately.
  • Reconcile the data before approval: the supplied segment figures imply approximately 32% overall retention, rather than 34%. Confirm whether they use different cohorts or definitions.
  • Track onboarding reliability alongside these KRs. Critical onboarding defects are release blockers, not a bug-count target.
  • Improve collaboration through genuine value—not forced invitations or changes that artificially inflate membership.

Candidate initiatives — not scored key results

Prioritize the checklist, invite-flow improvements, a small set of useful templates, and lifecycle emails according to their expected contribution to these outcomes. Use interviews and funnel analysis to choose and refine experiments.

Defer dark mode and AI summaries unless evidence shows they address a major onboarding barrier. Publish help content where research identifies a specific need.

Note to the squad

I replaced three broad objectives and a shipping checklist with one focused objective and measurable user outcomes. Launches, interviews, articles, and bug fixes are ways to make progress; completing them does not prove that onboarding improved.

Early team participation is our strongest available signal: teams with three or more people joining in week 1 have much higher week-4 retention. That is correlation, not proof of causation, so we’ll test whether our changes improve retention rather than assume invitations alone will do it.

I also narrowed the scope for our five-engineer, one-designer squad. NPS is too broad to diagnose this onboarding problem, while dark mode and AI summaries would compete with the core work.

Because OKR scores now affect bonuses, I want targets, scoring rules, guardrails, and shared ownership agreed before the quarter—not adjusted afterward. We should review the bonus linkage with leadership and HR so it does not reward metric manipulation or discourage honest experiments.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly58%
  • passAddresses the actual decision47%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty59%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable73%
  • passBuilds the key results on the data100%
  • passFlags the bonus link98%
  • passKey results are outcomes, not output98%
  • passFocuses on the big rock100%
  • passShows how the goals add up56%
Run
Run
#1
API response time
27 s
Submitted
5 Oct 2026

Needs more than a quick edit8 of 11 checks passed

The output fails to address the bonus-link risk properly and makes an unsupported claim about the biggest retention lever, requiring edits before use.

Output

To the Onboarding Squad:

Thanks for putting these together. I’ve rewritten our OKRs for leadership. Before I explain why, please note HR’s rule: our bonuses are now tied directly to our OKR scores. That means our OKRs must be about outcomes, not output, and they must tightly align with the company goal.

Here is what I changed:

  1. Dropped the "feature laundry list." Draft 1 listed outputs (checklists, dark mode, 50 bugs). If we ship all those and week 4 retention doesn't budge, leadership fails us—and our bonuses drop. We need to focus on what happens to the user, not what we build.
  2. Anchored on the Company Goal: Leadership wants week 4 team retention up from 34% to 40%. Our entire quarter must serve this.
  3. Used our Data: Our data shows a massive retention cliff: teams where 3+ people join in week one retain at 58%, versus 21% for everyone else. Right now, only 29% hit that 3+ threshold. Furthermore, the median time to a shared note is 3 days—too slow.

Here is the finalized version going to leadership on Friday:

*

Q3 Onboarding Squad OKRs

Objective: Get new teams collaborating faster so they stick around.

  • KR 1: Raise the share of new teams with 3+ active members in week 1 from 29% to 40%. (Directly attacks our biggest retention lever).
  • KR 2: Reduce the median time from sign-up to the first note opened by a second team member from 3 days to under 24 hours. (Accelerates "time-to-value" and teamwork).
  • KR 3: Raise the share of new teams still active in week 4 from 34% to 40%. (Our direct contribution to the company goal).

*

What this means for us: Features like the checklist, templates, AI summaries, and invite flow redesign are still our tools to hit these goals, but we’ll prioritize them based on whether they actually drive team invites and fast collaboration. Let’s crush this.

Grades and run details

Decision model 82 · LLM judge 8 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly76%
  • passAddresses the actual decision71%
  • passRespects explicit constraints56%
  • failIdentifies material uncertainty87%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable92%
  • passBuilds the key results on the data100%
  • partialFlags the bonus link93%
  • passKey results are outcomes, not output99%
  • passFocuses on the big rock100%
  • passShows how the goals add up79%
Run
Run
#1
Time to output
3 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Identifies material uncertaintyWrongRightWrong
GPT-6 Luna · API

The output does not name any specific unknowns that could change the OKRs or how they would be resolved.

GPT-6.1 Sol · API

It identifies the data discrepancy, the correlation vs. causation risk, and the need for feasibility review, and says how to resolve them.

Gemini 3.5 Flash-Lite · Gemini

The output does not name any unknowns that could change the decision or how they would be resolved.

Avoids unsupported claimsMixedRightWrong
GPT-6 Luna · API

The statement 'getting a note opened by someone else is a concrete sign of shared value' is presented as fact without being labelled as an assumption or hypothesis, and the evidence does not establish it.

GPT-6.1 Sol · API

Hypotheses and causes are clearly labelled (e.g., 'correlation, not proof of causation'), and no interpretations are presented as established fact.

Gemini 3.5 Flash-Lite · Gemini

The claim 'biggest retention lever' is presented as fact without support from the data, which only shows correlation.

Flags the bonus linkWrongRightWrong
GPT-6 Luna · API

The output mentions that OKR scores affect bonuses but does not flag the risk of safe targets or suggest separating bonuses from OKR scores.

GPT-6.1 Sol · API

The output flags the risk of bonus-linked OKRs leading to safe targets or manipulation, and suggests reviewing the linkage with leadership and HR.

Gemini 3.5 Flash-Lite · Gemini

The output notes the bonus-OKR link but does not suggest separating them or another concrete fix for the risk of safe targets.

All got right 8

Uses the supplied evidence correctlyRightRightRight
GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context.

GPT-6.1 Sol · API

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

All factual claims about the current situation are directly from the brief or supplied context.

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

The output commits to a clear set of rewritten OKRs and a note, which is the decision the brief asked for.

GPT-6.1 Sol · API

The output commits to a clear rewritten OKR set and a note explaining the changes, framed for the squad, and states that targets are proposed pending feasibility review.

Gemini 3.5 Flash-Lite · Gemini

The output commits to a clear set of rewritten OKRs and a note, addressing the request.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output is a rewritten OKR set with a note, well under 600 words, and respects the form and reader.

GPT-6.1 Sol · API

The output is a note to the squad with rewritten OKRs, well under 600 words, and respects the requested form and reader.

Gemini 3.5 Flash-Lite · Gemini

The output is a note to the squad with the OKRs, well under 600 words, respecting the form and length.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The deliverable is a complete rewritten OKR set and a note, within the word limit, and usable by the squad.

GPT-6.1 Sol · API

The deliverable is a complete, usable OKR rewrite and squad note, in the right form and within the word limit.

Gemini 3.5 Flash-Lite · Gemini

The note and OKRs are complete, within length, and usable by the squad with light edits.

Builds the key results on the dataRightRightRight
GPT-6 Luna · API

Key results 2 and 3 are built on the first-week joining data and the time to a shared note, with baselines from the brief and targets.

GPT-6.1 Sol · API

Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.

Gemini 3.5 Flash-Lite · Gemini

KR1 and KR2 are built on the first-week joining data and time to shared note, with baselines and targets.

Key results are outcomes, not outputRightRightRight
GPT-6 Luna · API

All key results are measurable changes in customer behaviour with baselines and targets; shipping work appears only as candidate initiatives.

GPT-6.1 Sol · API

All key results are measurable outcomes with baselines and targets; shipping work appears only as candidate initiatives.

Gemini 3.5 Flash-Lite · Gemini

All key results are measurable changes in customer behavior, not things to ship.

Focuses on the big rockRightRightRight
GPT-6 Luna · API

The output cuts the goals to one objective with three key results, drops NPS, dark mode, AI summaries, bug counts, interviews and help articles, and explains why.

GPT-6.1 Sol · API

The output reduces the draft to one objective with three KRs, explicitly drops dark mode, AI summaries, and NPS, and explains why.

Gemini 3.5 Flash-Lite · Gemini

The output cuts the goals to one objective with three KRs and explains what was dropped.

Shows how the goals add upRightRightRight
GPT-6 Luna · API

Each key result is linked to the company goal, with the data on three-person teams and shared notes showing how they contribute.

GPT-6.1 Sol · API

Every key result is linked to the company retention goal, with KR1 directly shared and KRs 2-3 as leading indicators, and the reasoning is explained.

Gemini 3.5 Flash-Lite · Gemini

Each KR is linked to the company goal with reasoning or data, and no KR serves no company goal.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini77.362.52None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review