Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Focuses on the big rock100% pass
    The output reduces the draft to one objective with three KRs, explicitly drops dark mode, AI summaries, and NPS, and explains why.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  3. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Fieldnote's onboarding squad. The squad drafted its OKRs for next quarter and asked you to tidy them up before they go to leadership on Friday. Rewrite them, and add a short note to the squad on what you changed and why. No more than 600 words in all. What we know is below.

What the model was given6 items: About Fieldnote, The company's goal for next quarter (set by leadership), The squad's draft, Data, A note from HR, The squad
About FieldnoteShared notes and docs for teams, sold self-serve. Teams sign up free and pay once they pass five members.
The company's goal for next quarter (set by leadership)Objective: make Fieldnote the place new teams get started. Key result: raise the share of new teams still active in week 4 from 34% to 40%.
The squad's draftObjective 1: Ship the new onboarding. Key results: launch an onboarding checklist; ship 5 templates; redesign the invite flow; launch onboarding emails. Objective 2: Improve quality. Key results: fix 50 bugs; get P1 bugs to zero. Objective 3: Delight users. Key results: raise NPS from 31 to 40; launch dark mode; ship AI summaries; run 10 user interviews; publish 4 help articles.
DataNew teams where three or more people join in the first week: 58% still active in week 4. Everyone else: 21%. Today 29% of new teams get three or more people in the first week. Median time from sign-up to the first note someone else opens: 3 days.
A note from HRFrom this year, each squad member's bonus multiplier is set by their squad's OKR score.
The squadFive engineers, one designer and you.
What a strong answer doesThe answer key the graders mark against

Rewrites the draft into one objective tied to the company's (new teams get started) with two or three outcome key results built from the data, such as raising the share of new teams with three or more people in week one from 29% to a stated target, and cutting the time to the first shared note from 3 days to a stated target, with the company's week-4 retention as the result they serve. The checklist, templates, invite redesign and emails become initiatives, not key results. Drops what doesn't serve the goal (dark mode, AI summaries) or moves it out of the OKRs; keeps bug work as a health measure rather than an objective. Flags that tying bonuses to OKR scores will push the squad to set safe targets, and suggests keeping them apart. Explains the changes to the squad briefly and warmly.

Critical failures (cap the score at 40)
  • Keeps the draft's key results essentially unchanged
Task

Mid-level · v1.0 · synthetic · B2B SaaS, self-serve, goals

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit7 of 11 checks passed

The output includes unsupported claims about NPS and bug incentives, and fails to properly address the bonus-link risk, so a PM would need to rework those parts before use.

Output

Here is the revised OKR draft along with the team note.

*

Proposed Q3 OKRs: Onboarding Squad

Objective 1: Accelerate new teams to their first collaborative moment * KR 1: Increase the share of new teams with 3+ members in Week 1 from 29% to 45%. * KR 2: Reduce median time from team creation to first collaborative note opened from 3 days to <24 hours. * KR 3: Contribute directly to the company goal by lifting new team Week 4 retention from 34% to 40%.

Objective 2: Deliver a seamless, zero-friction first-run experience * KR 1: Maintain zero open P1 bugs across the sign-up, invite, and first-note flows throughout the quarter. * KR 2: Increase invite-flow conversion (invites sent per new team) by 25%.

*

Note to the Squad

Team,

Thanks for pulling the initial draft together! Ahead of our review with leadership this Friday, I’ve refocused our OKRs. Here is what changed and why:

#### 1. We shifted from outputs to outcomes Shipping checklists, templates, and emails are solutions, not goals. If we ship all four and retention doesn’t budge, we haven’t succeeded. By framing our KRs around moving specific user behaviors (e.g., getting teams to collaborate faster), we give ourselves the flexibility to iterate, drop what isn’t working, and double down on what does.

#### 2. We’re anchoring on our biggest retention lever Our company goal is lifting Week 4 retention from 34% to 40%. Our data shows an undeniable path to get there: * Teams with 3+ members in Week 1 retain at 58% (vs. 21% for everyone else). * Currently, only 29% of teams hit that milestone. * It takes 3 full days for a teammate to open their first shared note.

Our primary job this quarter is simple: get teammates into the workspace and collaborating on day one. Every template, invite prompt, or checklist we build should serve that specific loop.

#### 3. We cut scope to protect our focus (and our bonuses) Now that HR ties our bonus multipliers directly to squad OKR completion, our OKRs must be focused, measurable, and within our sphere of influence: * Dropped AI summaries and Dark Mode: Both are cool features, but neither solves the core first-week team activation problem. They dilute our engineering capacity. * Dropped NPS (31 to 40): NPS is a lagging, noisy metric influenced by the entire product (pricing, bugs elsewhere, feature gaps), not just onboarding. * Refined bug targets: Fixing an arbitrary "50 bugs" incentivizes fixing trivial issues. Instead, we’re committing to zero critical bugs in our core onboarding and invitation paths.

Let’s chat through any questions at standup tomorrow so we’re fully aligned before Friday!

Claims the judge couldn’t find in the brief

  • NPS is a lagging, noisy metric influenced by the entire product (pricing, bugs elsewhere, feature gaps), not just onboarding.
  • Fixing an arbitrary '50 bugs' incentivizes fixing trivial issues.

Grades and run details

Decision model 77 · LLM judge 7 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly28%
  • passAddresses the actual decision63%
  • passRespects explicit constraints19%
  • failIdentifies material uncertainty65%
  • partialAvoids unsupported claims31%
  • passProduces the required deliverable65%
  • passBuilds the key results on the data100%
  • partialFlags the bonus link82%
  • passKey results are outcomes, not output40%
  • passFocuses on the big rock94%
  • partialShows how the goals add up81%
Run
Run
#1
API response time
20 s
Submitted
5 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 3

Identifies material uncertaintyWrong

The output does not name any specific unknowns that could change the decision or how they would be resolved.

Avoids unsupported claimsWrong

It states as fact that NPS is lagging/noisy and that the 50-bug target incentivizes trivial issues, without labelling these as hypotheses or interpretations.

Flags the bonus linkWrong

The output mentions the bonus link but does not name the risk of safe targets or suggest separating bonuses from OKR scores; it only says OKRs must be focused.

Mixed 1

Uses the supplied evidence correctlyMixed

The output presents unsupported claims about NPS being lagging/noisy and the 50-bug target incentivizing trivial issues, which are not in the brief or derivable from it.

Got right 7

Addresses the actual decisionRight

The output commits to a clear revised OKR set and explains the changes, framed for the squad.

Respects explicit constraintsRight

The output is a rewritten OKR draft with a note, under 600 words, and respects the requested form and reader.

Produces the required deliverableRight

The revised OKRs and note are complete, within the word limit, and usable by the squad with light edits.

Builds the key results on the dataRight

KR1 uses the 29% baseline and sets a target; KR2 uses the 3-day baseline and sets a target; both are built on the supplied data.

Key results are outcomes, not outputRight

All key results are measurable changes in user behavior or quality (e.g., share of teams, time, retention, bug count, conversion), not shipped features.

Focuses on the big rockRight

The output reduces the draft to two objectives with a few KRs, drops dark mode, AI summaries, and NPS, and explains why.

Shows how the goals add upRight

Each KR is linked to the company goal via the retention data, and the note explains how the 3+ member and time-to-collaboration metrics drive Week 4 retention.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini77.362.52None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review