Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty66% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly71% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM working for Hannah Iyer, Quarry's CPO. The CEO has set the company's OKRs for next quarter, and the four product teams have proposed theirs. Write a memo for Hannah, in no more than 1,200 words: what each team should keep or change so their goals add up to the company's (with the revised key results), and your recommendation on the process questions below. The pack below is everything we have. Not all of it matters equally.

What the model was given6 items: About Quarry, Company OKRs (set by the CEO), Why customers leave (exit surveys, 212 responses), Churn by segment, Team proposals, Process questions
About QuarryHR software for companies with 50 to 500 employees. 420 staff, 2,600 customers. The core HR product is sold per employee; Payroll is a paid add-on.
Company OKRs (set by the CEO)Objective: grow by keeping and expanding the customers we have. Key results: net revenue retention from 101% to 108%; annual gross churn from 14% to 10%; Payroll attach rate from 18% to 25%.
Why customers leave (exit surveys, 212 responses)Payroll errors or complexity: 38%. Price: 24%. Missing integrations: 19%. Moved to an all-in-one competitor: 11%. Other: 8%.
Churn by segmentCustomers using Payroll: 8% a year. Customers without it: 16%. Customers with three or more integrations connected: 6%.
Team proposalsCore HR. Objective: build the best HR platform. Key results: ship the org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. Payroll. Objective: make payroll effortless. Key results: cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. Integrations. Objective: connect everything. Key result: ship 12 new integrations. Growth. Objective: win more customers. Key results: new logos up 20%; trial conversion from 9% to 12%.
Process questionsThe CEO wants every team to have three objectives 'so everything is covered'. The CFO wants OKR scores to set PM bonuses from next quarter. Last year teams scored 0.95 on average. Planning last quarter took five of the quarter's thirteen weeks.
What a strong answer doesThe answer key the graders mark against

Rewrites each team's goals so they add up to the company's. Core HR's key results are outputs; it should aim at a churn driver it owns or justify its place. Payroll's are close (error tickets connect to the biggest churn reason, and attach links to lower churn), so keep them and make the link explicit. Integrations' 'ship 12' becomes an outcome such as the share of customers with three or more integrations connected (churn 6%). Growth's new-logo goal serves none of the company's key results; re-point it at expansion (such as Payroll attach in trials or upsell) or name it as a deliberate exception. Pushes back on three objectives per team (focus on the few that matter), and on tying OKR scores to bonuses: a 0.95 average already suggests safe targets, and pay would make it worse. Suggests a lighter, shorter planning process.

Critical failures (cap the score at 40)
  • Endorses tying OKR scores to PM bonuses without flagging the risk
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The memo is clear, actionable, and covers all required elements; a PM could use it with only minor wording adjustments.

Output

To: Hannah Iyer, CPO Subject: Next-quarter OKRs: focus teams on retention and expansion

The CEO’s objective is clear, but the team proposals do not yet add up to it: Growth is optimizing new logos, while Core HR and Integrations mostly measure shipping. I recommend one objective per team, with a small number of measurable results tied to the company outcomes. Keep delivery plans such as “ship 12 integrations” as initiatives, not proof of impact.

Team recommendations

Core HR — change the objective and replace output KRs. “Build the best HR platform” is too broad. Org-chart redesign, custom fields, and CSAT may be useful, but the pack does not show that shipping them will reduce churn or grow revenue. Prioritize work against the customer problems behind churn—including price/value, which needs a product and pricing response rather than a feature assumption.

Revised objective: Improve retention of core HR customers. - Reduce annualized gross churn for customers without Payroll from 16% to at most 10.7%. This is a first-pass allocation: if Payroll attach reaches 25%, Payroll churn stays at 8%, and the segment rates are customer-weighted, non-Payroll churn needs to be about 10.7% for overall churn to reach 10%. Validate definitions and denominators before locking this target. - Keep CSAT at or above 4.2 as a guardrail, rather than making a score increase the main success measure.

Treat the two proposed launches as candidate initiatives, to be prioritized only if evidence links them to retention or expansion.

Payroll — keep the strongest proposal; clarify the outcome. Payroll directly supports both the company’s attach and churn goals. Payroll customers churn at half the rate of non-Payroll customers, and payroll errors or complexity are the most common reported reason for leaving.

Revised objective: Make Payroll reliable and grow adoption among current customers. - Increase Payroll attach from 18% to 25%. - Reduce payroll-error tickets by 40%. - Keep annual Payroll customer churn at 8% or lower.

Define error-ticket counting consistently, and pair the reduction with a quality guardrail so the metric cannot improve by discouraging customers from reporting problems.

Integrations — keep the problem area; replace “ship 12” as the key result. The evidence supports this area: missing integrations account for 19% of exit-survey responses, and customers with three or more integrations have 6% annual churn. But shipping integrations does not show that customers adopt them or stay.

Revised objective: Reduce integration-related customer loss. - Reduce integrations’ share of exit-survey reasons for leaving from 19% to 14%; treat this as a directional diagnostic, since survey shares are noisy. - Increase the number of customers with three or more active integrations. Establish the baseline in week one and set a numeric quarter target before plans are finalized; the pack does not provide enough data to invent one. - Keep “12 integrations shipped” as a delivery commitment only where customer demand and expected adoption justify it.

Growth — change the objective and stop optimizing acquisition this quarter. New logos and trial conversion do not support the stated company objective. Reorient Growth toward expansion in the installed base, working with Payroll on customer targeting and adoption.

Revised objective: Grow revenue from existing customers. - Increase company NRR from 101% to 108%; Growth is accountable for the expansion motion, with other teams contributing to the shared result. - Set a numeric target, after establishing a baseline, for expansion revenue from existing customers excluding Payroll attach. This makes Growth’s contribution visible without double-counting Payroll’s attach KR.

Process recommendations

  • Do not require three objectives per team. Coverage is not a reason to dilute focus. Use one objective and a few KRs per team; capture other responsibilities as initiatives or guardrails. Make cross-team ownership explicit, with one directly accountable owner for each company KR.
  • Do not use OKR scores to set PM bonuses next quarter. Bonuses would encourage teams to choose safe targets, avoid shared work, or optimize the score instead of customer outcomes. An average score of 0.95 last year is already a reason to examine target-setting and scoring calibration—not to add financial stakes. Keep performance and compensation decisions separate from quarterly OKR scores.
  • Cut planning from five weeks to two. Five weeks consumes nearly 40% of the quarter. Timebox baseline checks and target-setting, then revisit assumptions during the quarter. Use the missing integration and expansion baselines to sharpen targets early, not to prolong planning.

Finally, treat the company’s annual churn and NRR measures as lagging outcomes: review leading indicators during the quarter, but keep the company outcomes as the test of whether the portfolio is working.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision68%
  • passRespects explicit constraints54%
  • passIdentifies material uncertainty65%
  • passAvoids unsupported claims30%
  • partialProduces the required deliverable52%
  • passAims the teams with the churn data94%
  • passReads the 0.95 average for what it is91%
  • passKey results are outcomes, not output70%
  • passFocuses on the big rock97%
  • passShows how the goals add up57%
Run
Run
#1
API response time
57 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is complete, evidence-based, and directly actionable; a PM could use it with only minor wording adjustments.

Output

To: Hannah Iyer, CPO From: Staff PM Subject: Next-quarter OKRs: one retention-and-expansion plan, not four roadmaps

Recommendation

Give each team one focused objective, with two or three measurable results. Redirect Growth toward existing customers, replace shipping commitments with customer outcomes, and retain Payroll’s strong alignment.

The evidence favors Payroll reliability and adoption, followed by integrations. Payroll problems account for 38% of surveyed exits; missing integrations account for 19%. Price, at 24%, also needs an explicit response—not necessarily a price cut.

These are priorities, not causal proof. Exit surveys may not represent all churned customers, and customers using Payroll or several integrations may differ systematically from others. We should test whether interventions improve retention rather than assume adoption alone causes it.

Proposed team OKRs

The additional targets below are proposed planning targets, not forecasts. Validate feasibility and missing baselines during the first planning week; do not spend the quarter negotiating them.

Core HR: Make the core product worth staying for

Keep: Improving customer satisfaction, but as supporting evidence rather than the primary definition of success.

Change: “Build the best HR platform” is too broad. Shipping the org-chart redesign and custom fields v2 does not demonstrate retention impact. Keep either initiative only if customer evidence connects it to material retention risk; otherwise defer it.

Revised key results 1. Reduce annual gross customer churn among non-Payroll customers from 16% to 11%. 2. Increase CSAT from 4.2 to 4.5, using the same survey method and a comparable customer population.

Core HR is the DRI for non-Payroll retention, working with Customer Success, Growth, and Integrations—not its sole contributor. Prioritize recurring problems among at-risk customers over broad redesigns. Use CSAT diagnostically: a satisfaction increase without better retention is not sufficient success.

Payroll: Make Payroll reliable enough to adopt and keep

Keep: The objective and both proposed results. This is the clearest connection to the company strategy.

Change: Normalize support tickets for usage, and explicitly measure retention among Payroll customers. Otherwise adoption growth can obscure whether reliability improved.

Revised key results 1. Reduce payroll-error tickets per 1,000 payroll runs by 40%, without reducing support accessibility or increasing severe payroll incidents. 2. Increase paid Payroll attach from 18% to 25%. 3. Reduce annual gross customer churn among Payroll customers from 8% to 7%.

Payroll owns the attach outcome and activation experience; Growth owns the commercial contribution. Treat attach as one shared metric, not two independently credited wins. Reliability must improve alongside adoption: expanding a troublesome product could increase churn.

Integrations: Connect the workflows that make customers stay

Keep: Investment in integrations.

Change: Replace “connect everything” and “ship 12 integrations.” Twelve low-demand connectors could produce no retention benefit. Select work using demand from existing customers, churn evidence, and activation friction; improving an existing connector may beat launching another.

Revised key results 1. Increase the share of the opening customer cohort with three or more actively used integrations by 10 percentage points. 2. Reduce integration-attributed churn events per 100 opening customers by 30%, versus a comparable prior-period baseline, with consistent reason coding.

Both are provisional targets requiring baseline validation in planning week one. Define active use as successful, recurring data exchange—not merely connecting an account. Track connector launches and reliability operationally, not as substitutes for these outcomes.

Growth: Expand retained customers, not just acquire new ones

Keep: Funnel measurement and experimentation capabilities.

Change: Remove new logos +20% and trial conversion 9%→12% from this quarter’s strategic OKRs. They may remain operating metrics, but they do not directly deliver the CEO’s existing-customer strategy.

Revised key results 1. Increase NRR from 101% to 108%, with Growth accountable for the commercial expansion plan and Hannah accountable for the cross-functional outcome. 2. Reduce price-attributed churn events per 100 opening customers by 25%, versus a comparable prior-period baseline.

Growth should focus on eligible existing customers’ Payroll discovery, purchase, and activation, alongside evidence-led pricing, packaging, and value communication. Do not meet the churn target through indiscriminate discounts: measure retained revenue after concessions and monitor contraction and margin. Payroll funnel conversion is a leading indicator; its baseline is not the new-customer trial conversion figure.

Make the results reconcile

The proposed segment targets provide a useful planning check: at 25% Payroll attach, 25% × 7% + 75% × 11% = 10% gross customer churn. This is a directional bridge, not a forecast: adoption timing, customer mix, and cohort definitions matter. The current segment figures imply 14.6%, rather than the reported 14%, so reconcile definitions before locking targets.

At an unchanged 2,600-customer denominator, 25% attach means 650 Payroll customers versus 468 today—182 net additions. The actual requirement changes with acquisition, churn, and add-on cancellations.

NRR needs a separate revenue bridge. Do not assume four points less customer churn yields four points more revenue retention. Finance and Growth should reconcile:

Opening recurring revenue − churned revenue − contraction + expansion = ending retained-customer revenue.

By planning close, show how expected Payroll expansion and revenue saved cover the seven-point NRR improvement, including employee-count changes and discounts. If they do not, the plan needs another expansion lever or a different resource allocation.

Use consistent annual churn and NRR definitions. A quarter cannot establish a full year’s causal retention effect; track the agreed company measure alongside quarterly cohort churn and leading indicators, without silently switching to an annualized quarterly rate.

Process decisions

Reject three objectives per team. Coverage belongs in roadmaps and operating dashboards. OKRs identify the few changes that matter most. One objective per team is sufficient here; maintain service, security, and reliability obligations as explicit guardrails.

Do not tie next quarter’s PM bonuses mechanically to OKR scores. That encourages conservative targets, denominator manipulation, and disputes over shared outcomes. Assess performance using results, decision quality, collaboration, and responsible execution. Distinguish committed from aspirational results before the quarter starts.

Last year’s 0.95 average warrants an audit, not praise or an automatic demand for lower scores. It could reflect excellent execution, easy targets, output-heavy goals, or inconsistent scoring. Review the underlying goals and evidence; do not impose an arbitrary score distribution.

Time-box planning to five working days. Five weeks consumed 38% of the quarter. Hannah should set priorities and trade-offs on day one; teams and Finance resolve baselines, dependencies, capacity, and the revenue bridge by day five. Use brief weekly metric checks and monthly cross-team reviews thereafter. Change goals only for explicit strategic reasons, never retrospectively to improve scores.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly42%
  • passAddresses the actual decision86%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable64%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is36%
  • passKey results are outcomes, not output94%
  • passFocuses on the big rock99%
  • passShows how the goals add up54%
Run
Run
#1
Time to output
82 s
Submitted
5 Oct 2026

Usable with a quick edit8 of 11 checks passed

The memo is actionable and well-structured; only a minor edit to change Integrations KR2 from a shipping deliverable to an outcome metric is needed, without redoing any reasoning.

Output

MEMORANDUM

TO: Hannah Iyer, Chief Product Officer FROM: Staff Product Manager DATE: October 24, 2023 SUBJECT: Q1 Product OKR Alignment & Planning Governance

---

Executive Summary

The CEO’s Q1 mandate is clear: grow through retention and expansion (NRR 101% → 108%, Gross Churn 14% → 10%, Payroll Attach 18% → 25%).

However, our product teams have largely proposed business-as-usual, output-driven goals. Most critically, Growth is running in the opposite direction (hunting new logos instead of retention/expansion), Integrations is measuring shipping volume instead of adoption, and Core HR has defaulted to feature delivery.

Below is an analysis of what each team must keep and change, their revised OKRs, and recommendations to resolve the three executive governance bottlenecks.

---

Part 1: Team-by-Team OKR Revisions

#### 1. Payroll * Keep: Focus on error reduction. Exit surveys show payroll errors/complexity cause 38% of all churn. Cutting error tickets is directly tied to saving revenue. * Change: Remove pure attach rate as Payroll’s sole downstream KR. Product cannot own GTM distribution alone; Payroll must focus on the product friction that unlocks adoption and prevents churn within the first 90 days. * Revised OKR: * Objective: Deliver a zero-defect payroll experience that drives expansion and retention. * KR 1: Reduce payroll-related support tickets / error reports by 40% (churn mitigation). * KR 2: Reduce time-to-first-run for newly activated payroll accounts from X days to Y days (activation velocity). * KR 3: Maintain monthly payroll run success rate at 99.5%+ across all active payroll customers.

#### 2. Integrations * Keep: Focus on ecosystem connectivity. Customers with 3+ integrations churn at just 6% (vs. 14% company average). * Change: Kill the output metric ("Ship 12 new integrations"). Integrations have zero ROI if unadopted. The team must optimize for depth of adoption—specifically pushing accounts to the 3+ threshold—and target the specific missing integrations driving the 19% churn response. * Revised OKR: * Objective: Deepen customer stickiness by embedding Quarry into the customer’s core toolstack. * KR 1: Increase % of total customers with 3+ active integrations from X% to Y% (driving toward the 6% churn cohort). * KR 2: Deliver the top 3 requested integrations identified in churn exit surveys (accounting/time-tracking). * KR 3: Increase weekly active sync volume / data exchange by 25%.

#### 3. Growth * Keep: Growth's experimentation velocity and funnel mechanics. * Change: Completely pivot away from top-of-funnel acquisition (new logos / trials). Acquisition is GTM-heavy and contradicts the company’s explicit retention/expansion strategy. Growth must be deployed against in-app cross-sell (Payroll attach) and per-seat expansion. * Revised OKR: * Objective: Accelerate revenue expansion across the installed customer base. * KR 1: Drive self-serve/assisted Payroll attach rate from 18% to 25% (CEO target). * KR 2: Increase in-app expansion qualified leads (PQLs) for Payroll passed to Sales by 35%. * KR 3: Increase customer employee invite/seat-activation rate from X% to Y% in accounts with >100 employees.

#### 4. Core HR * Keep: Focus on platform satisfaction (CSAT). * Change: Eliminate roadmap feature-shipping milestones ("Ship org-chart", "Ship custom fields v2") as Key Results. Features are inputs; the outcome is reducing platform-switching churn (11% leave for all-in-one competitors) and protecting the base. * Revised OKR: * Objective: Establish Core HR as the indispensable daily system of record for mid-market admins. * KR 1: Increase core platform CSAT among HR Admins from 4.2 to 4.5. * KR 2: Increase WAU/MAU engagement of company admins on core workflows (org management, employee records) from X% to Y%. * KR 3: Reduce voluntary logo churn of non-payroll accounts from 16% to 12%.

---

OKR Alignment Matrix

TeamFocus AreaPrimary Metric OwnedCompany OKR Impact
:---:---:---:---
PayrollDefect reduction & run reliabilityError tickets (-40%)Gross Churn (14% → 10%)
IntegrationsIntegration depth (3+ threshold)% accounts with 3+ integrationsGross Churn (14% → 10%)
GrowthCross-sell & expansion funnelsAttach rate (18% → 25%)NRR (101% → 108%)
Core HRPlatform utility & admin CSATAdmin CSAT (4.2 → 4.5) & base churnGross Churn (14% → 10%)

---

Part 2: Governance & Process Recommendations

#### 1. CEO Request: "Three Objectives per team to cover everything" * Recommendation: Decline firmly. * Rationale: A product team with 3 objectives across 4 teams creates 12 competing product priorities for 420 staff. When everything is covered, nothing is prioritized. Objectives are strategic shifts, not operational check-lists. * Counter-Proposal: Enforce one single Objective per team with 2–3 measurable, outcome-based Key Results. Day-to-day maintenance, technical debt, and minor bugs should be managed via standard SLAs and operational run-rate dashboards, not masqueraded as quarterly OKRs.

#### 2. CFO Request: Tie OKR scores to PM bonuses (Historical avg: 0.95) * Recommendation: Reject this proposal. * Rationale: A historical average of 0.95 is an alarming indicator of extreme sandbagging. Healthy OKR frameworks target an average score of 0.70 ("stretch goals"). If bonuses are linked to OKR achievement: 1. Teams will negotiate unchallenging, output-based metrics (e.g., "Ship custom fields" or "Ship 12 integrations") to guarantee 1.0 payouts. 2. Cross-team collaboration will break; teams will refuse shared accountability for complex metrics like Attach Rate or Churn. * Counter-Proposal: Separate evaluation from OKRs. Measure OKR outcomes as a lagging input into broader performance reviews, but assess PMs on holistic impact: strategic execution, product discovery quality, and business outcome delivery. Reset the target score expectations so that achieving 0.70 is celebrated as exceptional execution against ambitious goals.

#### 3. Planning Cadence: Reducing the 5-week overhead * Diagnosis: Spending 5 out of 13 weeks (38% of the quarter) planning is an unsustainable productivity drain. It signals an overly bottom-up, consensus-driven process with insufficient top-down constraints. * Prescription (2-Week Maximum Planning Cycle): * Week -2 (Day 1–3): Top-Down Mandate. Exec team delivers locked company OKRs and operational boundaries. * Week -2 (Day 4–7): Draft Team OKRs. Product squads define 1 Objective and 3 outcome-based KRs directly mapped to executive metrics. * Week -1 (Day 8–10): Cross-Dependency & CPO Review. Hannah and tech leads conduct a single cross-functional review to resolve dependencies (e.g., Growth ↔ Payroll attach). * Week -1 (Day 11–14): Final Approval & Lock. OKRs published before the start of the quarter; zero changes permitted thereafter.

---

Immediate Action Items for Hannah

  1. Schedule an Executive Alignment Session with the CEO: Present the unified 1-Objective model, demonstrating how 4 hyper-focused product goals cover Gross Churn, NRR, and Attach Rate far more effectively than 12 diluted objectives.
  2. Align with the CFO: Share the risk analysis regarding metric inflation (the 0.95 score reality) and present the alternative performance compensation model.
  3. Issue Directive to Product Leads: Reject Growth, Core HR, and Integrations' current drafts. Instruct them to submit revised KRs based on the outcome targets above by end of week.

Grades and run details

Decision model 59 · LLM judge 10 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly21%
  • passAddresses the actual decision70%
  • partialRespects explicit constraints26%
  • failIdentifies material uncertainty44%
  • partialAvoids unsupported claims44%
  • partialProduces the required deliverable62%
  • passAims the teams with the churn data99%
  • passReads the 0.95 average for what it is100%
  • partialKey results are outcomes, not output87%
  • passFocuses on the big rock89%
  • partialShows how the goals add up79%
Run
Run
#1
API response time
30 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context or simple arithmetic.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic.

Gemini 3.8 Flash · API

All factual claims about the current situation are directly from the brief or supplied context.

Identifies material uncertaintyRightRightWrong
GPT-6 Luna · API

It names missing baselines and definitional unknowns, and says to establish baselines in week one and validate definitions before locking targets.

GPT-6 Astra · ChatGPT

It names specific unknowns (survey representativeness, systematic differences, missing baselines) and says how to resolve them (test interventions, validate baselines in planning week one).

Gemini 3.8 Flash · API

The output does not name any specific unknowns that could change the decisions or say how they would be resolved.

Key results are outcomes, not outputRightRightWrong
GPT-6 Luna · API

Every key result is a measurable outcome (churn, attach, error tickets, CSAT guardrail, exit-survey share, 3+ integrations, NRR, expansion revenue); shipping appears only as initiatives.

GPT-6 Astra · ChatGPT

Every revised key result is a measurable change in customer or business behavior (churn rates, attach rates, NRR, CSAT, integration adoption), not a shipping commitment.

Gemini 3.8 Flash · API

Integrations KR2 ('Deliver the top 3 requested integrations') is a shipping deliverable, not an outcome metric.

All got right 8

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

The memo commits to clear revised key results for each team and explicit process recommendations, framed for Hannah.

GPT-6 Astra · ChatGPT

The memo commits to a clear set of recommendations (one objective per team, redirect Growth, keep Payroll alignment) and addresses the process questions unambiguously.

Gemini 3.8 Flash · API

The memo commits to clear recommendations for each team's OKRs and process questions, framed for Hannah.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output is a memo under 1,200 words, addresses the CPO, and respects the requested form.

GPT-6 Astra · ChatGPT

The output is a memo addressed to Hannah, stays within the 1,200-word limit, and respects the requested form.

Gemini 3.8 Flash · API

The output is a memo under 1,200 words, addressed to Hannah, and respects the requested form.

Avoids unsupported claimsRightRightRight
GPT-6 Luna · API

Interpretations and forecasts are clearly labelled or are directly supported by the evidence; no unsupported factual claims are presented as established.

GPT-6 Astra · ChatGPT

Hypotheses and causal interpretations are clearly labelled as such (e.g., 'priorities, not causal proof', 'we should test'), and no confident claim goes beyond the evidence.

Gemini 3.8 Flash · API

Interpretations and forecasts are presented as recommendations or rationale, not as established facts.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The memo provides revised team goals with key results and process recommendations, and is usable with light edits.

GPT-6 Astra · ChatGPT

The memo provides revised key results for each team, links them to company goals, and gives actionable process recommendations; a PM could use it with light edits.

Gemini 3.8 Flash · API

The memo includes revised OKRs, links to company goals, and process recommendations, and is usable as-is.

Aims the teams with the churn dataRightRightRight
GPT-6 Luna · API

It uses the churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, and Core HR at non-Payroll churn, with explicit links.

GPT-6 Astra · ChatGPT

It uses the churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, Core HR at non-Payroll churn, and Growth at price churn, with explicit links.

Gemini 3.8 Flash · API

The memo uses payroll errors (38%) for Payroll, 3+ integrations (6%) for Integrations, and 11% all-in-one for Core HR, with clear links.

Reads the 0.95 average for what it isRightRightRight
GPT-6 Luna · API

It flags the 0.95 average as a sign of safe targets and recommends against tying bonuses to scores to avoid further sandbagging.

GPT-6 Astra · ChatGPT

It recognizes the 0.95 average suggests safe targets, and explicitly advises against tying bonuses to OKR scores because it would encourage even safer targets.

Gemini 3.8 Flash · API

The memo flags the 0.95 average as sandbagging and argues that tying bonuses would worsen target-setting.

Focuses on the big rockRightRightRight
GPT-6 Luna · API

It cuts each team to one objective with a few KRs, drops shipping KRs, and explains what was dropped and why.

GPT-6 Astra · ChatGPT

It cuts each team to one objective with a few key results, drops new logos, trial conversion, and shipping goals, and explains why they were dropped.

Gemini 3.8 Flash · API

The memo cuts each team to one objective with 2-3 KRs, drops new logos and feature-shipping, and explains why.

Shows how the goals add upRightRightRight
GPT-6 Luna · API

Each team-level KR is linked to a company goal with reasoning or data, and the new-logo goal is flagged as not supporting the company objective and reoriented.

GPT-6 Astra · ChatGPT

Each key result is explicitly linked to a company goal (churn, NRR, attach) with reasoning, and the original Growth goals that served no company goal are flagged and replaced.

Gemini 3.8 Flash · API

Every team's primary metric is linked to a company OKR in the alignment matrix, and misaligned goals are flagged.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review