Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Focuses on the big rock100% pass
    The output reduces the draft to one objective with three KRs, explicitly drops dark mode, AI summaries, and NPS, and explains why.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  3. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty70% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly75% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM working for Hannah Iyer, Quarry's CPO. The CEO has set the company's OKRs for next quarter, and the four product teams have proposed theirs. Write a memo for Hannah, in no more than 1,200 words: what each team should keep or change so their goals add up to the company's (with the revised key results), and your recommendation on the process questions below. The pack below is everything we have. Not all of it matters equally.

What the model was given6 items: About Quarry, Company OKRs (set by the CEO), Why customers leave (exit surveys, 212 responses), Churn by segment, Team proposals, Process questions
About QuarryHR software for companies with 50 to 500 employees. 420 staff, 2,600 customers. The core HR product is sold per employee; Payroll is a paid add-on.
Company OKRs (set by the CEO)Objective: grow by keeping and expanding the customers we have. Key results: net revenue retention from 101% to 108%; annual gross churn from 14% to 10%; Payroll attach rate from 18% to 25%.
Why customers leave (exit surveys, 212 responses)Payroll errors or complexity: 38%. Price: 24%. Missing integrations: 19%. Moved to an all-in-one competitor: 11%. Other: 8%.
Churn by segmentCustomers using Payroll: 8% a year. Customers without it: 16%. Customers with three or more integrations connected: 6%.
Team proposalsCore HR. Objective: build the best HR platform. Key results: ship the org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. Payroll. Objective: make payroll effortless. Key results: cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. Integrations. Objective: connect everything. Key result: ship 12 new integrations. Growth. Objective: win more customers. Key results: new logos up 20%; trial conversion from 9% to 12%.
Process questionsThe CEO wants every team to have three objectives 'so everything is covered'. The CFO wants OKR scores to set PM bonuses from next quarter. Last year teams scored 0.95 on average. Planning last quarter took five of the quarter's thirteen weeks.
What a strong answer doesThe answer key the graders mark against

Rewrites each team's goals so they add up to the company's. Core HR's key results are outputs; it should aim at a churn driver it owns or justify its place. Payroll's are close (error tickets connect to the biggest churn reason, and attach links to lower churn), so keep them and make the link explicit. Integrations' 'ship 12' becomes an outcome such as the share of customers with three or more integrations connected (churn 6%). Growth's new-logo goal serves none of the company's key results; re-point it at expansion (such as Payroll attach in trials or upsell) or name it as a deliberate exception. Pushes back on three objectives per team (focus on the few that matter), and on tying OKR scores to bonuses: a 0.95 average already suggests safe targets, and pay would make it worse. Suggests a lighter, shorter planning process.

Critical failures (cap the score at 40)
  • Endorses tying OKR scores to PM bonuses without flagging the risk
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The memo is clear, evidence-based, and actionable; a PM could use it with only minor wording tweaks, without re-checking facts or redoing reasoning.

Output

To: Hannah Iyer, CPO From: Staff PM Subject: Align next quarter’s product goals to retention and expansion

Recommendation

Give each team one focused objective, supported by two or three measurable results. Replace acquisition and shipping commitments with customer outcomes that support 108% NRR, 10% annual gross churn, and 25% Payroll attach.

The strongest signals favor fixing Payroll and improving useful integration adoption. Payroll errors or complexity account for 38% of surveyed departures; missing integrations account for another 19%. Price, at 24%, also warrants attention to perceived value and commercial policy.

Payroll customers and customers with three or more integrations churn less, but these are correlations—not proof that attaching either product causes retention. We should improve adoption and reliability while testing their effect on retention, not force adoption to hit a number.

Revised team OKRs

Core HR: Help existing customers realize enough value to stay

Keep: CSAT improvement as a supporting indicator. Change: “Build the best HR platform” is too broad. Remove the org-chart redesign and custom fields v2 from the KRs; retain them as candidate initiatives only if evidence connects them to renewal risk or meaningful customer value.

Revised KRs: 1. Reduce annual gross churn from 14% to 10%, as a shared company result. Core HR is the coordinating product owner, not the sole contributor. 2. Raise CSAT from 4.2 to 4.5, using the same question, sampling approach, and customer population. 3. Improve retention at renewal by 10 percentage points in a pre-defined price/value-risk cohort, versus its baseline. Lock cohort criteria before intervention, include every eligible account, and establish the baseline in week one.

The third target is a proposed stretch target, not a forecast. Core should work with Customer Success and Finance on value realization and price objections. Product improvements alone may not address price sensitivity; unrestricted discounting would also undermine NRR.

Payroll: Make Payroll reliable and easy enough to adopt and keep

Keep: The objective’s intent and the focus on errors. Change: Make the error measure volume-adjusted, add a direct measure of complexity, and move accountability for attach to Growth. Payroll remains responsible for readiness and activation quality.

Revised KRs: 1. Reduce payroll-error tickets per 100 payroll runs by 40% against the previous-quarter baseline, with consistent ticket classification. 2. Reduce median customer time to complete a payroll run by 20%, comparing similar payroll complexity; establish the baseline in week one.

Track error severity and support-contact behavior alongside ticket volume. Fewer tickets are not success if errors persist or customers stop reporting them. Growth should not scale attach ahead of Payroll’s ability to deliver a dependable experience.

Integrations: Make the connections customers need dependable and useful

Keep: Investment in integrations. Change: Replace “connect everything” and “ship 12 integrations.” Twelve low-demand releases could satisfy the proposal without retaining anyone. Prioritize missing connections linked to renewal risk and demand among existing customers.

Revised KRs: 1. Increase by 20% the share of active customers using at least three healthy integrations, relative to a week-one baseline. “Using” must require successful recurring data exchange, not merely installation. 2. Reduce failed scheduled syncs per 1,000 sync attempts by 30%, versus the previous quarter.

Treat both numerical targets as planning proposals to validate against the baseline and capacity. Report retention for newly adopting customers against a comparable cohort. The observed 6% churn rate among customers with three or more integrations is a useful signal, not a guaranteed outcome for new adopters.

Growth: Expand revenue from existing customers through successful Payroll adoption

Keep: Experimentation and conversion discipline. Change: Replace “win more customers,” new-logo growth, and trial conversion. They support acquisition, not this quarter’s stated company objective. Necessary acquisition work can continue as business-as-usual; it should not dominate these OKRs.

Revised KRs: 1. Increase Payroll attach from 18% to 25%, measured on the same active-customer denominator as the company metric. 2. Increase NRR from 101% to 108%, as a shared company result, with Growth accountable for coordinating the expansion plan.

Growth owns targeting, commercial conversion, and the adoption funnel; Payroll owns product readiness. Track successful first payroll and subsequent usage so paid-but-unused attachments do not masquerade as progress. Evaluate incentives against margin, cancellations, and net retained revenue.

Make the goals add up financially

In week one, Finance and Analytics should produce a single retained-revenue bridge: opening recurring revenue, churn, contraction, expansion, and closing retained revenue. Size Payroll expansion and other expansion opportunities against the gap to 108% NRR.

We cannot infer that moving churn down four points and attach up seven points automatically achieves NRR. We lack account revenue, Payroll pricing, contraction, and the precise churn definition. Do not treat customer churn and revenue churn as interchangeable.

Preserve the company metrics’ existing definitions and reporting windows. Annual churn and NRR will respond slowly; review renewal outcomes, payroll quality, and adoption weekly without substituting those leading indicators for the company targets. Shared KRs should appear once on the executive scorecard, with named coordinating owners—not be counted repeatedly as independent contributions.

Process decisions

Reject three objectives per team. Coverage belongs in roadmaps and operating plans. A mandatory count encourages teams to invent goals and dilute attention. One objective per team is sufficient here.

Do not tie next-quarter PM bonuses mechanically to OKR scores. Teams share outcomes, annual metrics lag, and direct compensation linkage encourages sandbagging and metric manipulation. Use OKRs as evidence in a broader performance review covering impact, judgment, collaboration, and learning. Keep incentive decisions separate from stretch-goal scoring.

Investigate last year’s 0.95 average. It may indicate conservative targets, output-heavy KRs, or generous scoring—not exceptional performance. Calibrate committed versus stretch goals and review supporting evidence rather than imposing an arbitrary desired average score.

Time-box planning to two weeks. Five weeks consumed 38% of the quarter. Use week one for baselines, the revenue bridge, dependencies, and target feasibility; week two for trade-offs and Hannah’s approval. Review progress monthly and adjust initiatives as evidence changes, without quietly rewriting targets or definitions.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly45%
  • passAddresses the actual decision77%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty86%
  • passAvoids unsupported claims65%
  • passProduces the required deliverable65%
  • passAims the teams with the churn data95%
  • passReads the 0.95 average for what it is51%
  • passKey results are outcomes, not output91%
  • passFocuses on the big rock95%
  • passShows how the goals add up43%
Run
Run
#1
API response time
69 s
Submitted
5 Oct 2026

Usable with a quick edit8 of 11 checks passed

The memo is actionable and well-structured; only a minor edit to change Integrations KR2 from a shipping deliverable to an outcome metric is needed, without redoing any reasoning.

Output

MEMORANDUM

TO: Hannah Iyer, Chief Product Officer FROM: Staff Product Manager DATE: October 24, 2023 SUBJECT: Q1 Product OKR Alignment & Planning Governance

---

Executive Summary

The CEO’s Q1 mandate is clear: grow through retention and expansion (NRR 101% → 108%, Gross Churn 14% → 10%, Payroll Attach 18% → 25%).

However, our product teams have largely proposed business-as-usual, output-driven goals. Most critically, Growth is running in the opposite direction (hunting new logos instead of retention/expansion), Integrations is measuring shipping volume instead of adoption, and Core HR has defaulted to feature delivery.

Below is an analysis of what each team must keep and change, their revised OKRs, and recommendations to resolve the three executive governance bottlenecks.

---

Part 1: Team-by-Team OKR Revisions

#### 1. Payroll * Keep: Focus on error reduction. Exit surveys show payroll errors/complexity cause 38% of all churn. Cutting error tickets is directly tied to saving revenue. * Change: Remove pure attach rate as Payroll’s sole downstream KR. Product cannot own GTM distribution alone; Payroll must focus on the product friction that unlocks adoption and prevents churn within the first 90 days. * Revised OKR: * Objective: Deliver a zero-defect payroll experience that drives expansion and retention. * KR 1: Reduce payroll-related support tickets / error reports by 40% (churn mitigation). * KR 2: Reduce time-to-first-run for newly activated payroll accounts from X days to Y days (activation velocity). * KR 3: Maintain monthly payroll run success rate at 99.5%+ across all active payroll customers.

#### 2. Integrations * Keep: Focus on ecosystem connectivity. Customers with 3+ integrations churn at just 6% (vs. 14% company average). * Change: Kill the output metric ("Ship 12 new integrations"). Integrations have zero ROI if unadopted. The team must optimize for depth of adoption—specifically pushing accounts to the 3+ threshold—and target the specific missing integrations driving the 19% churn response. * Revised OKR: * Objective: Deepen customer stickiness by embedding Quarry into the customer’s core toolstack. * KR 1: Increase % of total customers with 3+ active integrations from X% to Y% (driving toward the 6% churn cohort). * KR 2: Deliver the top 3 requested integrations identified in churn exit surveys (accounting/time-tracking). * KR 3: Increase weekly active sync volume / data exchange by 25%.

#### 3. Growth * Keep: Growth's experimentation velocity and funnel mechanics. * Change: Completely pivot away from top-of-funnel acquisition (new logos / trials). Acquisition is GTM-heavy and contradicts the company’s explicit retention/expansion strategy. Growth must be deployed against in-app cross-sell (Payroll attach) and per-seat expansion. * Revised OKR: * Objective: Accelerate revenue expansion across the installed customer base. * KR 1: Drive self-serve/assisted Payroll attach rate from 18% to 25% (CEO target). * KR 2: Increase in-app expansion qualified leads (PQLs) for Payroll passed to Sales by 35%. * KR 3: Increase customer employee invite/seat-activation rate from X% to Y% in accounts with >100 employees.

#### 4. Core HR * Keep: Focus on platform satisfaction (CSAT). * Change: Eliminate roadmap feature-shipping milestones ("Ship org-chart", "Ship custom fields v2") as Key Results. Features are inputs; the outcome is reducing platform-switching churn (11% leave for all-in-one competitors) and protecting the base. * Revised OKR: * Objective: Establish Core HR as the indispensable daily system of record for mid-market admins. * KR 1: Increase core platform CSAT among HR Admins from 4.2 to 4.5. * KR 2: Increase WAU/MAU engagement of company admins on core workflows (org management, employee records) from X% to Y%. * KR 3: Reduce voluntary logo churn of non-payroll accounts from 16% to 12%.

---

OKR Alignment Matrix

TeamFocus AreaPrimary Metric OwnedCompany OKR Impact
:---:---:---:---
PayrollDefect reduction & run reliabilityError tickets (-40%)Gross Churn (14% → 10%)
IntegrationsIntegration depth (3+ threshold)% accounts with 3+ integrationsGross Churn (14% → 10%)
GrowthCross-sell & expansion funnelsAttach rate (18% → 25%)NRR (101% → 108%)
Core HRPlatform utility & admin CSATAdmin CSAT (4.2 → 4.5) & base churnGross Churn (14% → 10%)

---

Part 2: Governance & Process Recommendations

#### 1. CEO Request: "Three Objectives per team to cover everything" * Recommendation: Decline firmly. * Rationale: A product team with 3 objectives across 4 teams creates 12 competing product priorities for 420 staff. When everything is covered, nothing is prioritized. Objectives are strategic shifts, not operational check-lists. * Counter-Proposal: Enforce one single Objective per team with 2–3 measurable, outcome-based Key Results. Day-to-day maintenance, technical debt, and minor bugs should be managed via standard SLAs and operational run-rate dashboards, not masqueraded as quarterly OKRs.

#### 2. CFO Request: Tie OKR scores to PM bonuses (Historical avg: 0.95) * Recommendation: Reject this proposal. * Rationale: A historical average of 0.95 is an alarming indicator of extreme sandbagging. Healthy OKR frameworks target an average score of 0.70 ("stretch goals"). If bonuses are linked to OKR achievement: 1. Teams will negotiate unchallenging, output-based metrics (e.g., "Ship custom fields" or "Ship 12 integrations") to guarantee 1.0 payouts. 2. Cross-team collaboration will break; teams will refuse shared accountability for complex metrics like Attach Rate or Churn. * Counter-Proposal: Separate evaluation from OKRs. Measure OKR outcomes as a lagging input into broader performance reviews, but assess PMs on holistic impact: strategic execution, product discovery quality, and business outcome delivery. Reset the target score expectations so that achieving 0.70 is celebrated as exceptional execution against ambitious goals.

#### 3. Planning Cadence: Reducing the 5-week overhead * Diagnosis: Spending 5 out of 13 weeks (38% of the quarter) planning is an unsustainable productivity drain. It signals an overly bottom-up, consensus-driven process with insufficient top-down constraints. * Prescription (2-Week Maximum Planning Cycle): * Week -2 (Day 1–3): Top-Down Mandate. Exec team delivers locked company OKRs and operational boundaries. * Week -2 (Day 4–7): Draft Team OKRs. Product squads define 1 Objective and 3 outcome-based KRs directly mapped to executive metrics. * Week -1 (Day 8–10): Cross-Dependency & CPO Review. Hannah and tech leads conduct a single cross-functional review to resolve dependencies (e.g., Growth ↔ Payroll attach). * Week -1 (Day 11–14): Final Approval & Lock. OKRs published before the start of the quarter; zero changes permitted thereafter.

---

Immediate Action Items for Hannah

  1. Schedule an Executive Alignment Session with the CEO: Present the unified 1-Objective model, demonstrating how 4 hyper-focused product goals cover Gross Churn, NRR, and Attach Rate far more effectively than 12 diluted objectives.
  2. Align with the CFO: Share the risk analysis regarding metric inflation (the 0.95 score reality) and present the alternative performance compensation model.
  3. Issue Directive to Product Leads: Reject Growth, Core HR, and Integrations' current drafts. Instruct them to submit revised KRs based on the outcome targets above by end of week.

Grades and run details

Decision model 59 · LLM judge 10 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly21%
  • passAddresses the actual decision70%
  • partialRespects explicit constraints26%
  • failIdentifies material uncertainty44%
  • partialAvoids unsupported claims44%
  • partialProduces the required deliverable62%
  • passAims the teams with the churn data99%
  • passReads the 0.95 average for what it is100%
  • partialKey results are outcomes, not output87%
  • passFocuses on the big rock89%
  • partialShows how the goals add up79%
Run
Run
#1
API response time
30 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 11 checks passed

The memo is complete, evidence-based, and actionable; a PM could use it with only minor wording adjustments, without re-checking facts or redoing analysis.

Output

To: Hannah Iyer, CPO From: Staff PM Re: Q-next OKRs: making team goals add up, and the process questions

The short version

As proposed, the four teams' OKRs do not add up to the company's. Only Payroll's goals clearly support the CEO's objective. Integrations is measuring output rather than outcome. Core HR is pointed at its own roadmap. Growth is pointed at new customers, which the company OKRs deliberately leave out.

I recommend one objective per team, each tied to one or more of the three company KRs, with outcome-based key results. On process, I recommend no to three objectives per team, no to bonuses tied to OKR scores in their current form, and a two-week cap on planning.

How the company numbers can be reached

The company KRs are linked. Net revenue retention (NRR) is roughly 100% minus churned revenue plus net expansion.

  • Churn. Cutting gross churn from 14% to 10% is worth about 4 points of NRR.
  • Expansion. The other ~3 points have to come from expansion. Because Payroll is a paid add-on, attach (18% → 25%) is the main lever we control.
  • What finance should check. Whether 7 points of attach yields ~3 points of NRR depends on Payroll's price relative to core seats. That is worth checking before we commit to the 108% target.

Where churn comes from:

  • Exit reasons (212 surveys, about 364 churned customers a year): payroll errors or complexity 38%, price 24%, missing integrations 19%, moved to an all-in-one competitor 11%.
  • Payroll users churn at 8%, versus 16% for customers without Payroll.
  • Customers with three or more integrations churn at 6%.

Two caveats shape the recommendations:

  1. Attach alone barely moves churn. Moving attach from 18% to 25% at today's segment churn rates lowers blended churn by only about half a point (from roughly 14.6% to 14% on these figures). And that assumes the relationship is causal; engaged customers may simply buy Payroll and stay. Most of the 4-point churn reduction has to come from fixing payroll errors and from integrations.
  2. Payroll errors are the biggest single churn reason. They must be fixed before we push attach hard. Selling more Payroll while it still drives 38% of exits could raise churn rather than reduce it.

Team by team

Payroll: keep, with small changes - Objective: keep "Make payroll effortless". - Keep: cut payroll-error tickets by 40%, as a leading indicator. - Add: gross churn among Payroll customers from 8% to 6%. This is the outcome that matters, and tickets can fall without customers noticing. - Change ownership of attach: the 18% → 25% attach KR should be jointly owned with Growth. Payroll can make the product sellable; it does not run the cross-sell motion. - Sequencing: error fixes ship in the first half of the quarter, and the attach push ramps after.

Integrations: change the key result - Objective: change to "Make Quarry the hub customers can't unplug." - Replace "ship 12 integrations" with: share of customers with three or more integrations connected, from [baseline] to [baseline + 10 pts]. The baseline should be confirmed in week one. - Add: activation rate of new integrations within 60 days of launch. - How to choose integrations: prioritise by the systems named in exit surveys and churn-risk accounts, not by count. Four integrations customers actually connect beat twelve they don't. - Test the causal claim: run a controlled push to drive integration adoption in one cohort. The 6% churn figure may reflect selection rather than effect.

Core HR: refocus - Problem: org-chart redesign and custom fields v2 are deliverables, not results. "Build the best HR platform" has no line of sight to retention. CSAT is not a company KR. - Objective: change to "Give customers without Payroll a reason to stay." This is the 16%-churn segment, which is most of the base. - New KR 1: gross churn among non-Payroll customers from 16% to 13%. - New KR 2: a usage-based health score for at-risk accounts, with a target set once the baseline is known. - Existing projects: keep org chart and custom fields only if the team can show they address retention. For example, custom fields v2 could matter if integrations need them for data mapping. - Capacity: if the team can't show that link, I'd lend some of its engineering capacity to Payroll's error work for the quarter.

Growth: redirect - Problem: new logos and trial conversion don't count toward NRR, churn or attach. New customers are excluded from NRR by definition. - Objective: change to "Expand within our existing customers." - New KR 1: Payroll attach 18% → 25%, shared with Payroll, through in-product trials and cross-sell for existing customers. - New KR 2: conversion from Payroll trial to paid among existing customers, from [baseline] to [target]. - Your call before we tell Growth: if the CEO still wants new-logo growth, that is a fourth company KR and should be added explicitly. Otherwise Growth is working off-strategy.

A gap no product team owns

Price drives 24% of exits, and no team can own that through product work alone. Price is also often a polite stand-in for "not worth it", which the work above partly addresses. Packaging and pricing need an owner from commercial or finance. I suggest you raise this with the CEO and CFO rather than assign it to a product team.

Revised OKRs at a glance

TeamObjectiveKey resultsCompany KR served
PayrollMake payroll effortlessError tickets −40%; Payroll-customer churn 8% → 6%; attach 18% → 25% (shared)Churn, attach, NRR
IntegrationsMake Quarry the hub customers can't unplugCustomers with 3+ integrations +10 pts; 60-day activation of new integrationsChurn, NRR
Core HRGive non-Payroll customers a reason to stayNon-Payroll churn 16% → 13%; health-score targetChurn, NRR
GrowthExpand within our existing customersAttach 18% → 25% (shared); existing-customer Payroll trial → paidAttach, NRR

Process recommendations

1. Three objectives per team: recommend against. - The company itself has one objective. Twelve team objectives would spread effort thin and reproduce the problem this memo fixes: goals that cover everything and add up to nothing. - Proposal for the CEO: one objective per team, two or three KRs each, all tied to company KRs. - Meeting his real concern: "everything covered" belongs in a health-metrics dashboard (uptime, support SLAs, CSAT) that each team maintains. That separates running the business from changing it, and gives him the coverage he wants without diluting focus.

2. OKR scores setting PM bonuses: recommend against as proposed. - Last year's average score of 0.95 already suggests goals were set to be hit, not stretched. A healthy stretch average is roughly 0.6–0.7. - Paying on scores will push every team to sandbag, just as the company sets genuinely hard targets like churn from 14% to 10%. It also discourages the honest scoring we need to learn what works. - Alternatives for the CFO: - Tie a portion of PM bonuses to company KR results, shared across teams, which encourages collaboration on the shared attach and churn goals. - Base the individual portion on manager assessment of the quality of the bets and the execution, with OKR scores as one input rather than a formula. - If the CFO insists on a formula: at minimum, calibrate targets centrally so teams can't set their own bar.

3. Planning took 5 of 13 weeks: cap it at two, finished before the quarter starts. - Week −3: company OKRs final. This quarter's are already set. - Week −2: teams draft one objective and two or three KRs against them. This memo can be the template. - Week −1: one joint review with you to resolve shared ownership and gaps. Then lock. - During the quarter: baselines marked [TBD] get filled in during week one. They should not reopen the plan.

Decisions I need from you

  1. Agree the revised team OKRs above, especially redirecting Growth.
  2. Decide whether new-logo growth is in or out. If in, it goes to the CEO as an explicit company KR.
  3. Raise pricing ownership, and the process recommendations, with the CEO and CFO.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly9%
  • passAddresses the actual decision89%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims35%
  • partialProduces the required deliverable60%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is100%
  • passKey results are outcomes, not output31%
  • passFocuses on the big rock99%
  • passShows how the goals add up85%
Run
Run
#1
Time to output
54 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyRightMixedMixed
GPT-6.1 Sol · API

All statements about the current situation are taken directly from the brief or context, or follow from them; no invented facts.

Gemini 3.8 Flash · API

All factual claims about the current situation are directly from the brief or supplied context.

Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic; no invented numbers or data.

Identifies material uncertaintyRightWrongRight
GPT-6.1 Sol · API

It identifies correlation vs. causation, missing data, and the need to validate targets, and says how to resolve them (retained-revenue bridge, testing, baselines).

Gemini 3.8 Flash · API

The output does not name any specific unknowns that could change the decisions or say how they would be resolved.

Opus 5.5 · Claude

It names key unknowns (causal vs. selection effects for integrations and Payroll, NRR arithmetic, baselines) and suggests how to resolve them.

Key results are outcomes, not outputRightWrongRight
GPT-6.1 Sol · API

Every key result is a measurable change in customer or business behavior with a baseline and target; no shipping outputs appear as KRs.

Gemini 3.8 Flash · API

Integrations KR2 ('Deliver the top 3 requested integrations') is a shipping deliverable, not an outcome metric.

Opus 5.5 · Claude

Every revised key result is a measurable outcome (churn, attach, activation, health score) with baselines and targets; shipping work is treated as initiatives.

All got right 8

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The memo commits to clear revised OKRs and process recommendations for Hannah, and notes what would change the call (testing retention effects, validating baselines).

Gemini 3.8 Flash · API

The memo commits to clear recommendations for each team's OKRs and process questions, framed for Hannah.

Opus 5.5 · Claude

The memo commits to clear revised OKRs and process recommendations, framed for the CPO, and states what would change the call (e.g., CEO's stance on new logos).

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

The output is a memo to Hannah, under 1,200 words, and respects the requested form and reader.

Gemini 3.8 Flash · API

The output is a memo under 1,200 words, addressed to Hannah, and respects the requested form.

Opus 5.5 · Claude

The output is a memo addressed to Hannah Iyer, stays within the 1,200-word limit, and respects the requested form.

Avoids unsupported claimsRightRightRight
GPT-6.1 Sol · API

Interpretations are labelled as such (correlations, signals, may indicate), and confident claims are supported by the evidence.

Gemini 3.8 Flash · API

Interpretations and forecasts are presented as recommendations or rationale, not as established facts.

Opus 5.5 · Claude

Hypotheses and causal claims are clearly labelled as such (e.g., 'assumes the relationship is causal', 'may reflect selection rather than effect').

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

The memo provides complete revised KRs and process recommendations, within length, and is directly usable by Hannah.

Gemini 3.8 Flash · API

The memo includes revised OKRs, links to company goals, and process recommendations, and is usable as-is.

Opus 5.5 · Claude

The memo provides complete revised OKRs, process recommendations, and a clear ask; the CPO could act on it with minimal edits.

Aims the teams with the churn dataRightRightRight
GPT-6.1 Sol · API

Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.

Gemini 3.8 Flash · API

The memo uses payroll errors (38%) for Payroll, 3+ integrations (6%) for Integrations, and 11% all-in-one for Core HR, with clear links.

Opus 5.5 · Claude

It uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll customers, with explicit links.

Reads the 0.95 average for what it isRightRightRight
GPT-6.1 Sol · API

Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.

Gemini 3.8 Flash · API

The memo flags the 0.95 average as sandbagging and argues that tying bonuses would worsen target-setting.

Opus 5.5 · Claude

It flags that the 0.95 average suggests safe targets and argues that tying bonuses to scores would worsen sandbagging, with a clear recommendation against.

Focuses on the big rockRightRightRight
GPT-6.1 Sol · API

Cuts each team to one objective with a few KRs, drops shipping and acquisition goals, and explains why each was dropped.

Gemini 3.8 Flash · API

The memo cuts each team to one objective with 2-3 KRs, drops new logos and feature-shipping, and explains why.

Opus 5.5 · Claude

It cuts each team to one objective with a few KRs, explicitly drops non-contributing work (org chart, custom fields, new logos, shipping count), and explains why.

Shows how the goals add upRightRightRight
GPT-6.1 Sol · API

Every KR is linked to a company goal (churn, NRR, attach) with reasoning, and the new-logo goal is explicitly cut because it serves no company goal.

Gemini 3.8 Flash · API

Every team's primary metric is linked to a company OKR in the alignment matrix, and misaligned goals are flagged.

Opus 5.5 · Claude

The table and narrative link every team KR to a company KR (churn, attach, NRR), flag Growth's original new-logo goal as off-strategy, and note the unowned pricing gap.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini77.362.52None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review