Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty66% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly71% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM working for Hannah Iyer, Quarry's CPO. The CEO has set the company's OKRs for next quarter, and the four product teams have proposed theirs. Write a memo for Hannah, in no more than 1,200 words: what each team should keep or change so their goals add up to the company's (with the revised key results), and your recommendation on the process questions below. The pack below is everything we have. Not all of it matters equally.

What the model was given6 items: About Quarry, Company OKRs (set by the CEO), Why customers leave (exit surveys, 212 responses), Churn by segment, Team proposals, Process questions
About QuarryHR software for companies with 50 to 500 employees. 420 staff, 2,600 customers. The core HR product is sold per employee; Payroll is a paid add-on.
Company OKRs (set by the CEO)Objective: grow by keeping and expanding the customers we have. Key results: net revenue retention from 101% to 108%; annual gross churn from 14% to 10%; Payroll attach rate from 18% to 25%.
Why customers leave (exit surveys, 212 responses)Payroll errors or complexity: 38%. Price: 24%. Missing integrations: 19%. Moved to an all-in-one competitor: 11%. Other: 8%.
Churn by segmentCustomers using Payroll: 8% a year. Customers without it: 16%. Customers with three or more integrations connected: 6%.
Team proposalsCore HR. Objective: build the best HR platform. Key results: ship the org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. Payroll. Objective: make payroll effortless. Key results: cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. Integrations. Objective: connect everything. Key result: ship 12 new integrations. Growth. Objective: win more customers. Key results: new logos up 20%; trial conversion from 9% to 12%.
Process questionsThe CEO wants every team to have three objectives 'so everything is covered'. The CFO wants OKR scores to set PM bonuses from next quarter. Last year teams scored 0.95 on average. Planning last quarter took five of the quarter's thirteen weeks.
What a strong answer doesThe answer key the graders mark against

Rewrites each team's goals so they add up to the company's. Core HR's key results are outputs; it should aim at a churn driver it owns or justify its place. Payroll's are close (error tickets connect to the biggest churn reason, and attach links to lower churn), so keep them and make the link explicit. Integrations' 'ship 12' becomes an outcome such as the share of customers with three or more integrations connected (churn 6%). Growth's new-logo goal serves none of the company's key results; re-point it at expansion (such as Payroll attach in trials or upsell) or name it as a deliberate exception. Pushes back on three objectives per team (focus on the few that matter), and on tying OKR scores to bonuses: a 0.95 average already suggests safe targets, and pay would make it worse. Suggests a lighter, shorter planning process.

Critical failures (cap the score at 40)
  • Endorses tying OKR scores to PM bonuses without flagging the risk
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 11 checks passed

The memo is complete, evidence-based, and actionable; a PM could use it with only minor wording adjustments, without re-checking facts or redoing analysis.

Output

To: Hannah Iyer, CPO From: Staff PM Re: Q-next OKRs: making team goals add up, and the process questions

The short version

As proposed, the four teams' OKRs do not add up to the company's. Only Payroll's goals clearly support the CEO's objective. Integrations is measuring output rather than outcome. Core HR is pointed at its own roadmap. Growth is pointed at new customers, which the company OKRs deliberately leave out.

I recommend one objective per team, each tied to one or more of the three company KRs, with outcome-based key results. On process, I recommend no to three objectives per team, no to bonuses tied to OKR scores in their current form, and a two-week cap on planning.

How the company numbers can be reached

The company KRs are linked. Net revenue retention (NRR) is roughly 100% minus churned revenue plus net expansion.

  • Churn. Cutting gross churn from 14% to 10% is worth about 4 points of NRR.
  • Expansion. The other ~3 points have to come from expansion. Because Payroll is a paid add-on, attach (18% → 25%) is the main lever we control.
  • What finance should check. Whether 7 points of attach yields ~3 points of NRR depends on Payroll's price relative to core seats. That is worth checking before we commit to the 108% target.

Where churn comes from:

  • Exit reasons (212 surveys, about 364 churned customers a year): payroll errors or complexity 38%, price 24%, missing integrations 19%, moved to an all-in-one competitor 11%.
  • Payroll users churn at 8%, versus 16% for customers without Payroll.
  • Customers with three or more integrations churn at 6%.

Two caveats shape the recommendations:

  1. Attach alone barely moves churn. Moving attach from 18% to 25% at today's segment churn rates lowers blended churn by only about half a point (from roughly 14.6% to 14% on these figures). And that assumes the relationship is causal; engaged customers may simply buy Payroll and stay. Most of the 4-point churn reduction has to come from fixing payroll errors and from integrations.
  2. Payroll errors are the biggest single churn reason. They must be fixed before we push attach hard. Selling more Payroll while it still drives 38% of exits could raise churn rather than reduce it.

Team by team

Payroll: keep, with small changes - Objective: keep "Make payroll effortless". - Keep: cut payroll-error tickets by 40%, as a leading indicator. - Add: gross churn among Payroll customers from 8% to 6%. This is the outcome that matters, and tickets can fall without customers noticing. - Change ownership of attach: the 18% → 25% attach KR should be jointly owned with Growth. Payroll can make the product sellable; it does not run the cross-sell motion. - Sequencing: error fixes ship in the first half of the quarter, and the attach push ramps after.

Integrations: change the key result - Objective: change to "Make Quarry the hub customers can't unplug." - Replace "ship 12 integrations" with: share of customers with three or more integrations connected, from [baseline] to [baseline + 10 pts]. The baseline should be confirmed in week one. - Add: activation rate of new integrations within 60 days of launch. - How to choose integrations: prioritise by the systems named in exit surveys and churn-risk accounts, not by count. Four integrations customers actually connect beat twelve they don't. - Test the causal claim: run a controlled push to drive integration adoption in one cohort. The 6% churn figure may reflect selection rather than effect.

Core HR: refocus - Problem: org-chart redesign and custom fields v2 are deliverables, not results. "Build the best HR platform" has no line of sight to retention. CSAT is not a company KR. - Objective: change to "Give customers without Payroll a reason to stay." This is the 16%-churn segment, which is most of the base. - New KR 1: gross churn among non-Payroll customers from 16% to 13%. - New KR 2: a usage-based health score for at-risk accounts, with a target set once the baseline is known. - Existing projects: keep org chart and custom fields only if the team can show they address retention. For example, custom fields v2 could matter if integrations need them for data mapping. - Capacity: if the team can't show that link, I'd lend some of its engineering capacity to Payroll's error work for the quarter.

Growth: redirect - Problem: new logos and trial conversion don't count toward NRR, churn or attach. New customers are excluded from NRR by definition. - Objective: change to "Expand within our existing customers." - New KR 1: Payroll attach 18% → 25%, shared with Payroll, through in-product trials and cross-sell for existing customers. - New KR 2: conversion from Payroll trial to paid among existing customers, from [baseline] to [target]. - Your call before we tell Growth: if the CEO still wants new-logo growth, that is a fourth company KR and should be added explicitly. Otherwise Growth is working off-strategy.

A gap no product team owns

Price drives 24% of exits, and no team can own that through product work alone. Price is also often a polite stand-in for "not worth it", which the work above partly addresses. Packaging and pricing need an owner from commercial or finance. I suggest you raise this with the CEO and CFO rather than assign it to a product team.

Revised OKRs at a glance

TeamObjectiveKey resultsCompany KR served
PayrollMake payroll effortlessError tickets −40%; Payroll-customer churn 8% → 6%; attach 18% → 25% (shared)Churn, attach, NRR
IntegrationsMake Quarry the hub customers can't unplugCustomers with 3+ integrations +10 pts; 60-day activation of new integrationsChurn, NRR
Core HRGive non-Payroll customers a reason to stayNon-Payroll churn 16% → 13%; health-score targetChurn, NRR
GrowthExpand within our existing customersAttach 18% → 25% (shared); existing-customer Payroll trial → paidAttach, NRR

Process recommendations

1. Three objectives per team: recommend against. - The company itself has one objective. Twelve team objectives would spread effort thin and reproduce the problem this memo fixes: goals that cover everything and add up to nothing. - Proposal for the CEO: one objective per team, two or three KRs each, all tied to company KRs. - Meeting his real concern: "everything covered" belongs in a health-metrics dashboard (uptime, support SLAs, CSAT) that each team maintains. That separates running the business from changing it, and gives him the coverage he wants without diluting focus.

2. OKR scores setting PM bonuses: recommend against as proposed. - Last year's average score of 0.95 already suggests goals were set to be hit, not stretched. A healthy stretch average is roughly 0.6–0.7. - Paying on scores will push every team to sandbag, just as the company sets genuinely hard targets like churn from 14% to 10%. It also discourages the honest scoring we need to learn what works. - Alternatives for the CFO: - Tie a portion of PM bonuses to company KR results, shared across teams, which encourages collaboration on the shared attach and churn goals. - Base the individual portion on manager assessment of the quality of the bets and the execution, with OKR scores as one input rather than a formula. - If the CFO insists on a formula: at minimum, calibrate targets centrally so teams can't set their own bar.

3. Planning took 5 of 13 weeks: cap it at two, finished before the quarter starts. - Week −3: company OKRs final. This quarter's are already set. - Week −2: teams draft one objective and two or three KRs against them. This memo can be the template. - Week −1: one joint review with you to resolve shared ownership and gaps. Then lock. - During the quarter: baselines marked [TBD] get filled in during week one. They should not reopen the plan.

Decisions I need from you

  1. Agree the revised team OKRs above, especially redirecting Growth.
  2. Decide whether new-logo growth is in or out. If in, it goes to the CEO as an explicit company KR.
  3. Raise pricing ownership, and the process recommendations, with the CEO and CFO.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly9%
  • passAddresses the actual decision89%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims35%
  • partialProduces the required deliverable60%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is100%
  • passKey results are outcomes, not output31%
  • passFocuses on the big rock99%
  • passShows how the goals add up85%
Run
Run
#1
Time to output
54 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is complete, evidence-based, and directly actionable; a PM could use it with only minor wording adjustments.

Output

To: Hannah Iyer, CPO From: Staff PM Subject: Next-quarter OKRs: one retention-and-expansion plan, not four roadmaps

Recommendation

Give each team one focused objective, with two or three measurable results. Redirect Growth toward existing customers, replace shipping commitments with customer outcomes, and retain Payroll’s strong alignment.

The evidence favors Payroll reliability and adoption, followed by integrations. Payroll problems account for 38% of surveyed exits; missing integrations account for 19%. Price, at 24%, also needs an explicit response—not necessarily a price cut.

These are priorities, not causal proof. Exit surveys may not represent all churned customers, and customers using Payroll or several integrations may differ systematically from others. We should test whether interventions improve retention rather than assume adoption alone causes it.

Proposed team OKRs

The additional targets below are proposed planning targets, not forecasts. Validate feasibility and missing baselines during the first planning week; do not spend the quarter negotiating them.

Core HR: Make the core product worth staying for

Keep: Improving customer satisfaction, but as supporting evidence rather than the primary definition of success.

Change: “Build the best HR platform” is too broad. Shipping the org-chart redesign and custom fields v2 does not demonstrate retention impact. Keep either initiative only if customer evidence connects it to material retention risk; otherwise defer it.

Revised key results 1. Reduce annual gross customer churn among non-Payroll customers from 16% to 11%. 2. Increase CSAT from 4.2 to 4.5, using the same survey method and a comparable customer population.

Core HR is the DRI for non-Payroll retention, working with Customer Success, Growth, and Integrations—not its sole contributor. Prioritize recurring problems among at-risk customers over broad redesigns. Use CSAT diagnostically: a satisfaction increase without better retention is not sufficient success.

Payroll: Make Payroll reliable enough to adopt and keep

Keep: The objective and both proposed results. This is the clearest connection to the company strategy.

Change: Normalize support tickets for usage, and explicitly measure retention among Payroll customers. Otherwise adoption growth can obscure whether reliability improved.

Revised key results 1. Reduce payroll-error tickets per 1,000 payroll runs by 40%, without reducing support accessibility or increasing severe payroll incidents. 2. Increase paid Payroll attach from 18% to 25%. 3. Reduce annual gross customer churn among Payroll customers from 8% to 7%.

Payroll owns the attach outcome and activation experience; Growth owns the commercial contribution. Treat attach as one shared metric, not two independently credited wins. Reliability must improve alongside adoption: expanding a troublesome product could increase churn.

Integrations: Connect the workflows that make customers stay

Keep: Investment in integrations.

Change: Replace “connect everything” and “ship 12 integrations.” Twelve low-demand connectors could produce no retention benefit. Select work using demand from existing customers, churn evidence, and activation friction; improving an existing connector may beat launching another.

Revised key results 1. Increase the share of the opening customer cohort with three or more actively used integrations by 10 percentage points. 2. Reduce integration-attributed churn events per 100 opening customers by 30%, versus a comparable prior-period baseline, with consistent reason coding.

Both are provisional targets requiring baseline validation in planning week one. Define active use as successful, recurring data exchange—not merely connecting an account. Track connector launches and reliability operationally, not as substitutes for these outcomes.

Growth: Expand retained customers, not just acquire new ones

Keep: Funnel measurement and experimentation capabilities.

Change: Remove new logos +20% and trial conversion 9%→12% from this quarter’s strategic OKRs. They may remain operating metrics, but they do not directly deliver the CEO’s existing-customer strategy.

Revised key results 1. Increase NRR from 101% to 108%, with Growth accountable for the commercial expansion plan and Hannah accountable for the cross-functional outcome. 2. Reduce price-attributed churn events per 100 opening customers by 25%, versus a comparable prior-period baseline.

Growth should focus on eligible existing customers’ Payroll discovery, purchase, and activation, alongside evidence-led pricing, packaging, and value communication. Do not meet the churn target through indiscriminate discounts: measure retained revenue after concessions and monitor contraction and margin. Payroll funnel conversion is a leading indicator; its baseline is not the new-customer trial conversion figure.

Make the results reconcile

The proposed segment targets provide a useful planning check: at 25% Payroll attach, 25% × 7% + 75% × 11% = 10% gross customer churn. This is a directional bridge, not a forecast: adoption timing, customer mix, and cohort definitions matter. The current segment figures imply 14.6%, rather than the reported 14%, so reconcile definitions before locking targets.

At an unchanged 2,600-customer denominator, 25% attach means 650 Payroll customers versus 468 today—182 net additions. The actual requirement changes with acquisition, churn, and add-on cancellations.

NRR needs a separate revenue bridge. Do not assume four points less customer churn yields four points more revenue retention. Finance and Growth should reconcile:

Opening recurring revenue − churned revenue − contraction + expansion = ending retained-customer revenue.

By planning close, show how expected Payroll expansion and revenue saved cover the seven-point NRR improvement, including employee-count changes and discounts. If they do not, the plan needs another expansion lever or a different resource allocation.

Use consistent annual churn and NRR definitions. A quarter cannot establish a full year’s causal retention effect; track the agreed company measure alongside quarterly cohort churn and leading indicators, without silently switching to an annualized quarterly rate.

Process decisions

Reject three objectives per team. Coverage belongs in roadmaps and operating dashboards. OKRs identify the few changes that matter most. One objective per team is sufficient here; maintain service, security, and reliability obligations as explicit guardrails.

Do not tie next quarter’s PM bonuses mechanically to OKR scores. That encourages conservative targets, denominator manipulation, and disputes over shared outcomes. Assess performance using results, decision quality, collaboration, and responsible execution. Distinguish committed from aspirational results before the quarter starts.

Last year’s 0.95 average warrants an audit, not praise or an automatic demand for lower scores. It could reflect excellent execution, easy targets, output-heavy goals, or inconsistent scoring. Review the underlying goals and evidence; do not impose an arbitrary score distribution.

Time-box planning to five working days. Five weeks consumed 38% of the quarter. Hannah should set priorities and trade-offs on day one; teams and Finance resolve baselines, dependencies, capacity, and the revenue bridge by day five. Use brief weekly metric checks and monthly cross-team reviews thereafter. Change goals only for explicit strategic reasons, never retrospectively to improve scores.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly42%
  • passAddresses the actual decision86%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable64%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is36%
  • passKey results are outcomes, not output94%
  • passFocuses on the big rock99%
  • passShows how the goals add up54%
Run
Run
#1
Time to output
82 s
Submitted
5 Oct 2026

Needs more than a quick edit7 of 11 checks passed

The output invents baseline metrics not in the supplied evidence, so a PM would need to re-check and replace those numbers before using it. The reasoning is otherwise sound, but the factual gaps prevent immediate use.

Output

MEMORANDUM

TO: Hannah Iyer, Chief Product Officer FROM: Staff Product Manager DATE: October 24, 2023 SUBJECT: Q1 OKR Review & Planning Process Recommendations

---

Part 1: Strategic Alignment of Team OKRs

The CEO’s company-level OKR is clear: Grow by keeping and expanding the customers we have.

Our current metrics show the path forward clearly. Customers without Payroll churn at 16% annually, while those with it churn at 8%. Integrations act as a powerful retention anchor (churn drops to 6% for those with $\ge 3$ integrations). Exit surveys pinpoint our vulnerabilities: 38% leave due to payroll errors/complexity, and 19% due to missing integrations.

To hit a Net Revenue Retention (NRR) of 108%, gross churn of 10%, and a Payroll attach rate of 25%, our product teams cannot operate as independent feature factories. Their proposals currently suffer from output-focus (shipping lists of features) rather than outcome-focus.

Below is my assessment of what each team must keep or change, along with revised Key Results (KRs).

---

#### 1. Core HR * Current Proposal: Obj: Build the best HR platform. KRs: Ship org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. * Critique: "Build the best HR platform" is a vague umbrella. Org charts and custom fields are outputs, not outcomes. They do not directly map to NRR, churn reduction, or Payroll attach. * What to Keep: The focus on CSAT. Core HR is our baseline; low satisfaction bleeds into overall retention. * What to Change: Core HR must directly support the retention of non-Payroll customers (who churn at 16%) and lay the groundwork for Payroll discovery. * Revised OKRs: * Objective: Make Core HR indispensable to mid-market operations. * KR1: Raise Core HR CSAT from 4.2 to 4.5. * KR2: Increase 30-day post-signup feature adoption for accounts without Payroll from 40% to 65%. (Drives engagement to lower baseline churn). * KR3: Increase Payroll feature discovery interactions (clicks on Payroll upgrade prompts in Core HR) by 50%. (Supports company attach-rate goal).

---

#### 2. Payroll * Current Proposal: Obj: Make payroll effortless. KRs: Cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. * Critique: Excellent start. It directly addresses our #1 churn driver (38% leave due to payroll errors) and the company attach-rate goal. * What to Keep: Both KRs are exceptional. They tie directly to business metrics. * What to Change: The team needs a third KR focused on expansion revenue or reducing the complexity cited in exit surveys, rather than just error ticket volume. * Revised OKRs: * Objective: Make payroll effortless, reliable, and irresistible. * KR1: Cut payroll-error support tickets by 40%. * KR2: Increase Payroll attach rate from 18% to 25%. * KR3: Reduce time-to-run-payroll (median minutes per run) from 18 minutes to 10 minutes. (Directly attacks "complexity" cited in 38% of churned exit surveys).

---

#### 3. Integrations * Current Proposal: Obj: Connect everything. KR: Ship 12 new integrations. * Critique: "Ship 12 integrations" is a classic output trap. We could ship 12 unused integrations and see zero impact on churn. Customers with $\ge 3$ integrations churn at just 6%; 19% leave due to missing integrations. * What to Keep: The expansion of our ecosystem footprint. * What to Change: Pivot from counting shipped integrations to measuring active connection rates and their downstream impact on customer retention. * Revised OKRs: * Objective: Embed Quarry into our customers' existing tech stacks. * KR1: Increase the percentage of customers with 3+ active integrations connected from 22% to 35%. * KR2: Drive integration adoption in the top 5 most-requested categories (Accounting, Time-Tracking) to 40% of our customer base. * KR3: Reduce churn among customers missing key integrations by 30%.

---

#### 4. Growth * Current Proposal: Obj: Win more customers. KRs: New logos up 20%; trial conversion from 9% to 12%. * Critique: The company OKR explicitly shifts focus from pure acquisition to keeping and expanding what we have. A 20% increase in new logos conflicts with our need to fix retention and payroll friction. Furthermore, acquisition velocity without solved payroll/integration churn just feeds a leaky bucket. * What to Keep: Trial conversion optimization (efficient acquisition). * What to Change: Drop the new logo quota. Shift focus to acquiring the right logos (companies predisposed to adopting Payroll) and improving land-and-expand motions. * Revised OKRs: * Objective: Acquire and land high-LTV customers efficiently. * KR1: Increase free-to-paid trial conversion rate from 9% to 12%. * KR2: Increase trials starting with Payroll bundled or selected from 15% to 25%. * KR3: Reduce Day-90 logo churn for new cohorts from 5% to 3%.

---

Part 2: Recommendations on Process Questions

#### 1. The CEO’s request for three objectives per team * Recommendation: Push back firmly. * Rationale: Requiring three objectives per team (across four teams = 12 objectives) creates cognitive overload and diffuses focus. It forces teams to invent tertiary goals just to fill quota, leading to the exact output-driven KRs we are trying to eliminate. One sharp, outcome-oriented objective per team—tightly wound to the company’s NRR, gross churn, and attach-rate goals—is far more effective. If the CEO insists on structural symmetry, permit sub-objectives, but ensure each team has only one primary North Star objective for the quarter.

#### 2. The CFO’s proposal to use OKR scores for PM bonuses * Recommendation: Strongly advise against this. * Rationale: Tying OKR scores directly to compensation induces the "sandbagging effect." PMs will lobby for hockey-stick inputs disguised as KRs or set artificially low baselines to guarantee a 1.0 score (as evidenced by last year's suspicious 0.95 average). Furthermore, product outcomes depend on engineering execution, market conditions, and sales—factors outside a PM's absolute control. * Alternative: Use OKRs strictly for learning, strategic alignment, and performance calibration (what habits and rigor were applied), separating them entirely from bonus calculations. If bonuses must be tied to metrics, tie them to lagging company-level outcomes (like NRR or ARR), not team-level OKR self-assessments.

#### 3. Planning timeline (last quarter took 5 of 13 weeks) * Recommendation: Institute a strict 2-week timebox for Q1 planning, utilizing a Top-Down/Bottom-Up framework. * Rationale: Spending nearly 40% of a quarter on planning is a massive waste of R&D capital. The paralysis last quarter came from teams guessing what mattered because company goals weren't translated early enough. * Proposed Rules for Next Quarter: * Week -2 (led by CPO/CEO): Company OKRs locked and communicated (done). * Week 1 (led by PMs): Draft team OKRs against company constraints in 5 days. No multi-round negotiations. * Week 2 (led by CPO/Finance): Final cross-functional review, dependency mapping, and sign-off. * Pencils down by the end of Week 2. Any missing alignment is resolved by you (the CPO) rather than endless consensus-seeking.

Claims the judge couldn’t find in the brief

  • Core HR is our baseline.
  • The paralysis last quarter came from teams guessing what mattered because company goals weren't translated early enough.
  • Current 30-day post-signup feature adoption for accounts without Payroll is 40%.
  • Current time-to-run-payroll median is 18 minutes.
  • Current Day-90 logo churn for new cohorts is 5%.
  • Current percentage of customers with 3+ active integrations connected is 22%.
  • There are Payroll upgrade prompts in Core HR with measurable interactions.

Grades and run details

Decision model 73 · LLM judge 7 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly85%
  • passAddresses the actual decision56%
  • passRespects explicit constraints36%
  • failIdentifies material uncertainty68%
  • partialAvoids unsupported claims59%
  • passProduces the required deliverable69%
  • passAims the teams with the churn data97%
  • passReads the 0.95 average for what it is99%
  • passKey results are outcomes, not output35%
  • passFocuses on the big rock51%
  • partialShows how the goals add up30%
Run
Run
#1
Time to output
9 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyMixedRightWrong
Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic; no invented numbers or data.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic.

Gemini 3.5 Flash-Lite · Gemini

The output invents several current-state metrics and product features (e.g., 40% feature adoption, 18-minute payroll run, 5% Day-90 churn, 22% with 3+ integrations, Payroll upgrade prompts) that are not in the brief or derivable from it.

Addresses the actual decisionRightRightMixed
Opus 5.5 · Claude

The memo commits to clear revised OKRs and process recommendations, framed for the CPO, and states what would change the call (e.g., CEO's stance on new logos).

GPT-6 Astra · ChatGPT

The memo commits to a clear set of recommendations (one objective per team, redirect Growth, keep Payroll alignment) and addresses the process questions unambiguously.

Gemini 3.5 Flash-Lite · Gemini

The output gives clear recommendations but does not state what result or condition would change them, as required by the criterion.

Identifies material uncertaintyRightRightWrong
Opus 5.5 · Claude

It names key unknowns (causal vs. selection effects for integrations and Payroll, NRR arithmetic, baselines) and suggests how to resolve them.

GPT-6 Astra · ChatGPT

It names specific unknowns (survey representativeness, systematic differences, missing baselines) and says how to resolve them (test interventions, validate baselines in planning week one).

Gemini 3.5 Flash-Lite · Gemini

The output does not name any unknowns that could change the decision or say how they would be resolved.

Avoids unsupported claimsRightRightWrong
Opus 5.5 · Claude

Hypotheses and causal claims are clearly labelled as such (e.g., 'assumes the relationship is causal', 'may reflect selection rather than effect').

GPT-6 Astra · ChatGPT

Hypotheses and causal interpretations are clearly labelled as such (e.g., 'priorities, not causal proof', 'we should test'), and no confident claim goes beyond the evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents interpretations (e.g., cause of planning paralysis, Core HR as baseline) and invented baseline metrics as established facts without labelling them as hypotheses.

All got right 7

Respects explicit constraintsRightRightRight
Opus 5.5 · Claude

The output is a memo addressed to Hannah Iyer, stays within the 1,200-word limit, and respects the requested form.

GPT-6 Astra · ChatGPT

The output is a memo addressed to Hannah, stays within the 1,200-word limit, and respects the requested form.

Gemini 3.5 Flash-Lite · Gemini

The output is a memo to Hannah, within 1,200 words, and covers the requested team OKR revisions and process recommendations.

Produces the required deliverableRightRightRight
Opus 5.5 · Claude

The memo provides complete revised OKRs, process recommendations, and a clear ask; the CPO could act on it with minimal edits.

GPT-6 Astra · ChatGPT

The memo provides revised key results for each team, links them to company goals, and gives actionable process recommendations; a PM could use it with light edits.

Gemini 3.5 Flash-Lite · Gemini

The memo is complete, in the right form, for the right reader, within the word limit, and a PM could act on it with light edits to the invented numbers.

Aims the teams with the churn dataRightRightRight
Opus 5.5 · Claude

It uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll customers, with explicit links.

GPT-6 Astra · ChatGPT

It uses the churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, Core HR at non-Payroll churn, and Growth at price churn, with explicit links.

Gemini 3.5 Flash-Lite · Gemini

It uses the churn reasons and churn-by-segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll retention, with explicit links.

Reads the 0.95 average for what it isRightRightRight
Opus 5.5 · Claude

It flags that the 0.95 average suggests safe targets and argues that tying bonuses to scores would worsen sandbagging, with a clear recommendation against.

GPT-6 Astra · ChatGPT

It recognizes the 0.95 average suggests safe targets, and explicitly advises against tying bonuses to OKR scores because it would encourage even safer targets.

Gemini 3.5 Flash-Lite · Gemini

It calls the 0.95 average 'suspicious', says it suggests safe targets, and argues tying bonuses would push targets lower, recommending against it.

Key results are outcomes, not outputRightRightRight
Opus 5.5 · Claude

Every revised key result is a measurable outcome (churn, attach, activation, health score) with baselines and targets; shipping work is treated as initiatives.

GPT-6 Astra · ChatGPT

Every revised key result is a measurable change in customer or business behavior (churn rates, attach rates, NRR, CSAT, integration adoption), not a shipping commitment.

Gemini 3.5 Flash-Lite · Gemini

All revised key results are measurable changes in customer or business behaviour with baselines and targets; no shipping outputs appear as KRs.

Focuses on the big rockRightRightRight
Opus 5.5 · Claude

It cuts each team to one objective with a few KRs, explicitly drops non-contributing work (org chart, custom fields, new logos, shipping count), and explains why.

GPT-6 Astra · ChatGPT

It cuts each team to one objective with a few key results, drops new logos, trial conversion, and shipping goals, and explains why they were dropped.

Gemini 3.5 Flash-Lite · Gemini

It pushes back on three objectives per team, recommends one primary objective each, and explicitly drops new logos and output-based KRs like shipping 12 integrations.

Shows how the goals add upRightRightRight
Opus 5.5 · Claude

The table and narrative link every team KR to a company KR (churn, attach, NRR), flag Growth's original new-logo goal as off-strategy, and note the unowned pricing gap.

GPT-6 Astra · ChatGPT

Each key result is explicitly linked to a company goal (churn, NRR, attach) with reasoning, and the original Growth goals that served no company goal are flagged and replaced.

Gemini 3.5 Flash-Lite · Gemini

Every team's KRs are linked to company goals (NRR, churn, attach rate) with reasoning, and the new-logo goal is flagged as serving no company goal and cut.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review