Tasks / Metrics & Experimentation

Draft OKRs

Can the model turn a team's wish list into a few outcome-based OKRs that add up to the company's goals?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Builds the key results on the data100% pass
    Key results 2 and 3 are built directly on the first-week joining data and time-to-shared-note data, with baselines from the brief and targets.
    GPT-6.1 Sol · API · Eleven key results and a bonus
  2. Aims the teams with the churn data100% pass
    Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.
    GPT-6.1 Sol · API · The goals that don't add up
  3. Reads the 0.95 average for what it is100% pass
    Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.
    GPT-6.1 Sol · API · The goals that don't add up

Where it slips

  1. Flags the bonus link61% pass
    It names the risk of safe targets but only proposes asking leadership how a 70% score will be read, not a concrete fix like separating bonuses from OKR scores.
    Opus 5.5 · Claude · Eleven key results and a bonus
  2. Identifies material uncertainty66% pass
    The output does not name any specific unknowns that could change the OKRs or how they would be resolved.
    GPT-6 Luna · API · Eleven key results and a bonus
  3. Uses the supplied evidence correctly71% pass
    Claims that the gap is correlation and that the safe play is to pick easy KRs are not supported by the brief or arithmetic.
    Sonnet 5.5 · API · Eleven key results and a bonus

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Key results are outcomes, not output

    Is every key result a measurable change in customer or business behaviour, rather than something the team will ship or do?

    Passes when Every key result names a metric with a baseline and a target; shipping work appears only as initiatives that serve them.

  2. Focuses on the big rock

    Does the output cut the goals down to the few that matter most, rather than covering everything the team could do?

    Passes when Few objectives, each with a small number of key results, and says what was dropped and why.

  3. Shows how the goals add up

    Does the output show how each team-level key result contributes to the company-level goals, and flag any that don't?

    Passes when Every key result is linked to a company goal, with the reasoning or data that connects them, and any that serve no company goal are flagged or cut.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM working for Hannah Iyer, Quarry's CPO. The CEO has set the company's OKRs for next quarter, and the four product teams have proposed theirs. Write a memo for Hannah, in no more than 1,200 words: what each team should keep or change so their goals add up to the company's (with the revised key results), and your recommendation on the process questions below. The pack below is everything we have. Not all of it matters equally.

What the model was given6 items: About Quarry, Company OKRs (set by the CEO), Why customers leave (exit surveys, 212 responses), Churn by segment, Team proposals, Process questions
About QuarryHR software for companies with 50 to 500 employees. 420 staff, 2,600 customers. The core HR product is sold per employee; Payroll is a paid add-on.
Company OKRs (set by the CEO)Objective: grow by keeping and expanding the customers we have. Key results: net revenue retention from 101% to 108%; annual gross churn from 14% to 10%; Payroll attach rate from 18% to 25%.
Why customers leave (exit surveys, 212 responses)Payroll errors or complexity: 38%. Price: 24%. Missing integrations: 19%. Moved to an all-in-one competitor: 11%. Other: 8%.
Churn by segmentCustomers using Payroll: 8% a year. Customers without it: 16%. Customers with three or more integrations connected: 6%.
Team proposalsCore HR. Objective: build the best HR platform. Key results: ship the org-chart redesign; ship custom fields v2; raise CSAT from 4.2 to 4.5. Payroll. Objective: make payroll effortless. Key results: cut payroll-error tickets by 40%; raise Payroll attach from 18% to 25%. Integrations. Objective: connect everything. Key result: ship 12 new integrations. Growth. Objective: win more customers. Key results: new logos up 20%; trial conversion from 9% to 12%.
Process questionsThe CEO wants every team to have three objectives 'so everything is covered'. The CFO wants OKR scores to set PM bonuses from next quarter. Last year teams scored 0.95 on average. Planning last quarter took five of the quarter's thirteen weeks.
What a strong answer doesThe answer key the graders mark against

Rewrites each team's goals so they add up to the company's. Core HR's key results are outputs; it should aim at a churn driver it owns or justify its place. Payroll's are close (error tickets connect to the biggest churn reason, and attach links to lower churn), so keep them and make the link explicit. Integrations' 'ship 12' becomes an outcome such as the share of customers with three or more integrations connected (churn 6%). Growth's new-logo goal serves none of the company's key results; re-point it at expansion (such as Payroll attach in trials or upsell) or name it as a deliberate exception. Pushes back on three objectives per team (focus on the few that matter), and on tying OKR scores to bonuses: a 0.95 average already suggests safe targets, and pay would make it worse. Suggests a lighter, shorter planning process.

Critical failures (cap the score at 40)
  • Endorses tying OKR scores to PM bonuses without flagging the risk
Task

Staff-level · v1.0 · synthetic · B2B SaaS, goals, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The memo is complete, evidence-based, and directly actionable; a PM could use it with only minor wording adjustments.

Output

To: Hannah Iyer, CPO From: Staff PM Subject: Next-quarter OKRs: one retention-and-expansion plan, not four roadmaps

Recommendation

Give each team one focused objective, with two or three measurable results. Redirect Growth toward existing customers, replace shipping commitments with customer outcomes, and retain Payroll’s strong alignment.

The evidence favors Payroll reliability and adoption, followed by integrations. Payroll problems account for 38% of surveyed exits; missing integrations account for 19%. Price, at 24%, also needs an explicit response—not necessarily a price cut.

These are priorities, not causal proof. Exit surveys may not represent all churned customers, and customers using Payroll or several integrations may differ systematically from others. We should test whether interventions improve retention rather than assume adoption alone causes it.

Proposed team OKRs

The additional targets below are proposed planning targets, not forecasts. Validate feasibility and missing baselines during the first planning week; do not spend the quarter negotiating them.

Core HR: Make the core product worth staying for

Keep: Improving customer satisfaction, but as supporting evidence rather than the primary definition of success.

Change: “Build the best HR platform” is too broad. Shipping the org-chart redesign and custom fields v2 does not demonstrate retention impact. Keep either initiative only if customer evidence connects it to material retention risk; otherwise defer it.

Revised key results 1. Reduce annual gross customer churn among non-Payroll customers from 16% to 11%. 2. Increase CSAT from 4.2 to 4.5, using the same survey method and a comparable customer population.

Core HR is the DRI for non-Payroll retention, working with Customer Success, Growth, and Integrations—not its sole contributor. Prioritize recurring problems among at-risk customers over broad redesigns. Use CSAT diagnostically: a satisfaction increase without better retention is not sufficient success.

Payroll: Make Payroll reliable enough to adopt and keep

Keep: The objective and both proposed results. This is the clearest connection to the company strategy.

Change: Normalize support tickets for usage, and explicitly measure retention among Payroll customers. Otherwise adoption growth can obscure whether reliability improved.

Revised key results 1. Reduce payroll-error tickets per 1,000 payroll runs by 40%, without reducing support accessibility or increasing severe payroll incidents. 2. Increase paid Payroll attach from 18% to 25%. 3. Reduce annual gross customer churn among Payroll customers from 8% to 7%.

Payroll owns the attach outcome and activation experience; Growth owns the commercial contribution. Treat attach as one shared metric, not two independently credited wins. Reliability must improve alongside adoption: expanding a troublesome product could increase churn.

Integrations: Connect the workflows that make customers stay

Keep: Investment in integrations.

Change: Replace “connect everything” and “ship 12 integrations.” Twelve low-demand connectors could produce no retention benefit. Select work using demand from existing customers, churn evidence, and activation friction; improving an existing connector may beat launching another.

Revised key results 1. Increase the share of the opening customer cohort with three or more actively used integrations by 10 percentage points. 2. Reduce integration-attributed churn events per 100 opening customers by 30%, versus a comparable prior-period baseline, with consistent reason coding.

Both are provisional targets requiring baseline validation in planning week one. Define active use as successful, recurring data exchange—not merely connecting an account. Track connector launches and reliability operationally, not as substitutes for these outcomes.

Growth: Expand retained customers, not just acquire new ones

Keep: Funnel measurement and experimentation capabilities.

Change: Remove new logos +20% and trial conversion 9%→12% from this quarter’s strategic OKRs. They may remain operating metrics, but they do not directly deliver the CEO’s existing-customer strategy.

Revised key results 1. Increase NRR from 101% to 108%, with Growth accountable for the commercial expansion plan and Hannah accountable for the cross-functional outcome. 2. Reduce price-attributed churn events per 100 opening customers by 25%, versus a comparable prior-period baseline.

Growth should focus on eligible existing customers’ Payroll discovery, purchase, and activation, alongside evidence-led pricing, packaging, and value communication. Do not meet the churn target through indiscriminate discounts: measure retained revenue after concessions and monitor contraction and margin. Payroll funnel conversion is a leading indicator; its baseline is not the new-customer trial conversion figure.

Make the results reconcile

The proposed segment targets provide a useful planning check: at 25% Payroll attach, 25% × 7% + 75% × 11% = 10% gross customer churn. This is a directional bridge, not a forecast: adoption timing, customer mix, and cohort definitions matter. The current segment figures imply 14.6%, rather than the reported 14%, so reconcile definitions before locking targets.

At an unchanged 2,600-customer denominator, 25% attach means 650 Payroll customers versus 468 today—182 net additions. The actual requirement changes with acquisition, churn, and add-on cancellations.

NRR needs a separate revenue bridge. Do not assume four points less customer churn yields four points more revenue retention. Finance and Growth should reconcile:

Opening recurring revenue − churned revenue − contraction + expansion = ending retained-customer revenue.

By planning close, show how expected Payroll expansion and revenue saved cover the seven-point NRR improvement, including employee-count changes and discounts. If they do not, the plan needs another expansion lever or a different resource allocation.

Use consistent annual churn and NRR definitions. A quarter cannot establish a full year’s causal retention effect; track the agreed company measure alongside quarterly cohort churn and leading indicators, without silently switching to an annualized quarterly rate.

Process decisions

Reject three objectives per team. Coverage belongs in roadmaps and operating dashboards. OKRs identify the few changes that matter most. One objective per team is sufficient here; maintain service, security, and reliability obligations as explicit guardrails.

Do not tie next quarter’s PM bonuses mechanically to OKR scores. That encourages conservative targets, denominator manipulation, and disputes over shared outcomes. Assess performance using results, decision quality, collaboration, and responsible execution. Distinguish committed from aspirational results before the quarter starts.

Last year’s 0.95 average warrants an audit, not praise or an automatic demand for lower scores. It could reflect excellent execution, easy targets, output-heavy goals, or inconsistent scoring. Review the underlying goals and evidence; do not impose an arbitrary score distribution.

Time-box planning to five working days. Five weeks consumed 38% of the quarter. Hannah should set priorities and trade-offs on day one; teams and Finance resolve baselines, dependencies, capacity, and the revenue bridge by day five. Use brief weekly metric checks and monthly cross-team reviews thereafter. Change goals only for explicit strategic reasons, never retrospectively to improve scores.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly42%
  • passAddresses the actual decision86%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable64%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is36%
  • passKey results are outcomes, not output94%
  • passFocuses on the big rock99%
  • passShows how the goals add up54%
Run
Run
#1
Time to output
82 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is clear, evidence-based, and actionable; a PM could use it with only minor wording tweaks, without re-checking facts or redoing reasoning.

Output

To: Hannah Iyer, CPO From: Staff PM Subject: Align next quarter’s product goals to retention and expansion

Recommendation

Give each team one focused objective, supported by two or three measurable results. Replace acquisition and shipping commitments with customer outcomes that support 108% NRR, 10% annual gross churn, and 25% Payroll attach.

The strongest signals favor fixing Payroll and improving useful integration adoption. Payroll errors or complexity account for 38% of surveyed departures; missing integrations account for another 19%. Price, at 24%, also warrants attention to perceived value and commercial policy.

Payroll customers and customers with three or more integrations churn less, but these are correlations—not proof that attaching either product causes retention. We should improve adoption and reliability while testing their effect on retention, not force adoption to hit a number.

Revised team OKRs

Core HR: Help existing customers realize enough value to stay

Keep: CSAT improvement as a supporting indicator. Change: “Build the best HR platform” is too broad. Remove the org-chart redesign and custom fields v2 from the KRs; retain them as candidate initiatives only if evidence connects them to renewal risk or meaningful customer value.

Revised KRs: 1. Reduce annual gross churn from 14% to 10%, as a shared company result. Core HR is the coordinating product owner, not the sole contributor. 2. Raise CSAT from 4.2 to 4.5, using the same question, sampling approach, and customer population. 3. Improve retention at renewal by 10 percentage points in a pre-defined price/value-risk cohort, versus its baseline. Lock cohort criteria before intervention, include every eligible account, and establish the baseline in week one.

The third target is a proposed stretch target, not a forecast. Core should work with Customer Success and Finance on value realization and price objections. Product improvements alone may not address price sensitivity; unrestricted discounting would also undermine NRR.

Payroll: Make Payroll reliable and easy enough to adopt and keep

Keep: The objective’s intent and the focus on errors. Change: Make the error measure volume-adjusted, add a direct measure of complexity, and move accountability for attach to Growth. Payroll remains responsible for readiness and activation quality.

Revised KRs: 1. Reduce payroll-error tickets per 100 payroll runs by 40% against the previous-quarter baseline, with consistent ticket classification. 2. Reduce median customer time to complete a payroll run by 20%, comparing similar payroll complexity; establish the baseline in week one.

Track error severity and support-contact behavior alongside ticket volume. Fewer tickets are not success if errors persist or customers stop reporting them. Growth should not scale attach ahead of Payroll’s ability to deliver a dependable experience.

Integrations: Make the connections customers need dependable and useful

Keep: Investment in integrations. Change: Replace “connect everything” and “ship 12 integrations.” Twelve low-demand releases could satisfy the proposal without retaining anyone. Prioritize missing connections linked to renewal risk and demand among existing customers.

Revised KRs: 1. Increase by 20% the share of active customers using at least three healthy integrations, relative to a week-one baseline. “Using” must require successful recurring data exchange, not merely installation. 2. Reduce failed scheduled syncs per 1,000 sync attempts by 30%, versus the previous quarter.

Treat both numerical targets as planning proposals to validate against the baseline and capacity. Report retention for newly adopting customers against a comparable cohort. The observed 6% churn rate among customers with three or more integrations is a useful signal, not a guaranteed outcome for new adopters.

Growth: Expand revenue from existing customers through successful Payroll adoption

Keep: Experimentation and conversion discipline. Change: Replace “win more customers,” new-logo growth, and trial conversion. They support acquisition, not this quarter’s stated company objective. Necessary acquisition work can continue as business-as-usual; it should not dominate these OKRs.

Revised KRs: 1. Increase Payroll attach from 18% to 25%, measured on the same active-customer denominator as the company metric. 2. Increase NRR from 101% to 108%, as a shared company result, with Growth accountable for coordinating the expansion plan.

Growth owns targeting, commercial conversion, and the adoption funnel; Payroll owns product readiness. Track successful first payroll and subsequent usage so paid-but-unused attachments do not masquerade as progress. Evaluate incentives against margin, cancellations, and net retained revenue.

Make the goals add up financially

In week one, Finance and Analytics should produce a single retained-revenue bridge: opening recurring revenue, churn, contraction, expansion, and closing retained revenue. Size Payroll expansion and other expansion opportunities against the gap to 108% NRR.

We cannot infer that moving churn down four points and attach up seven points automatically achieves NRR. We lack account revenue, Payroll pricing, contraction, and the precise churn definition. Do not treat customer churn and revenue churn as interchangeable.

Preserve the company metrics’ existing definitions and reporting windows. Annual churn and NRR will respond slowly; review renewal outcomes, payroll quality, and adoption weekly without substituting those leading indicators for the company targets. Shared KRs should appear once on the executive scorecard, with named coordinating owners—not be counted repeatedly as independent contributions.

Process decisions

Reject three objectives per team. Coverage belongs in roadmaps and operating plans. A mandatory count encourages teams to invent goals and dilute attention. One objective per team is sufficient here.

Do not tie next-quarter PM bonuses mechanically to OKR scores. Teams share outcomes, annual metrics lag, and direct compensation linkage encourages sandbagging and metric manipulation. Use OKRs as evidence in a broader performance review covering impact, judgment, collaboration, and learning. Keep incentive decisions separate from stretch-goal scoring.

Investigate last year’s 0.95 average. It may indicate conservative targets, output-heavy KRs, or generous scoring—not exceptional performance. Calibrate committed versus stretch goals and review supporting evidence rather than imposing an arbitrary desired average score.

Time-box planning to two weeks. Five weeks consumed 38% of the quarter. Use week one for baselines, the revenue bridge, dependencies, and target feasibility; week two for trade-offs and Hannah’s approval. Review progress monthly and adjust initiatives as evidence changes, without quietly rewriting targets or definitions.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly45%
  • passAddresses the actual decision77%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty86%
  • passAvoids unsupported claims65%
  • passProduces the required deliverable65%
  • passAims the teams with the churn data95%
  • passReads the 0.95 average for what it is51%
  • passKey results are outcomes, not output91%
  • passFocuses on the big rock95%
  • passShows how the goals add up43%
Run
Run
#1
API response time
69 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 11 checks passed

The memo is complete, well-structured, and actionable; a PM could use it with minor edits to fill in baselines, which the memo already flags as needed.

Output

MEMO To: Hannah Iyer, CPO From: Staff PM Re: Aligning team OKRs to the company OKRs, and the process questions

Bottom line

Only Payroll's proposal is close to the company objective (keep and expand existing customers). Core HR and Integrations propose output goals (features and integrations shipped). Growth targets new logos, which move none of the CEO's three key results. I recommend one objective per team, no mechanical link between OKR scores and bonuses, and planning capped at two weeks.

The logic behind the changes

Applying the exit-survey shares to 14% annual churn gives roughly:

  • Payroll errors/complexity: 5.3 points
  • Price: 3.4 points
  • Missing integrations: 2.7 points
  • All-in-one competitor: 1.5 points
  • Other: 1.1 points

Getting to 10% means removing about 4 points.

Payroll attach alone won't get there. Moving attach from 18% to 25% is about 180 more customers. Payroll customers churn at 8% against 16% for the rest, so even if the gap were entirely causal, attach is worth about 0.6 points of churn. Much of the gap is probably selection, because engaged customers adopt more. The same caution applies to the 6% churn among customers with three or more integrations.

Attach could also backfire. Payroll errors are the top exit reason, so pushing a flawed product onto more customers risks creating more of the churn we are trying to prevent.

The bulk of the churn reduction has to come from fixing Payroll quality and integration gaps. I suggest this split, as a planning assumption rather than a forecast:

Exit poolOwnerTargetChurn points
Payroll errors/complexityPayrollcut ~40%~2.1
Missing integrationsIntegrationscut ~33%~0.9
PriceGrowthcut ~15%~0.5
All-in-one competitorCore HRcut ~33%~0.5

Together these come to about 4 points, which is 14% to 10%.

The exit data has limits. We have 212 responses against roughly 360 annual churners, and "payroll complexity" may include non-users who left for a payroll-capable competitor. Week one should confirm the cuts by segment.

Team by team

Payroll: keep most of it, tighten it. - Objective: Make payroll reliable enough that customers stay and others add it. - KR1: Payroll-error tickets per 1,000 pay runs down 40%. Normalise the metric. A raw ticket count will rise as attach grows, and it can also be gamed by making tickets harder to file. - KR2: Payroll-error/complexity-cited exits down about 40% (about 2 points of company churn). - KR3: Payroll attach from 18% to 25%, with a guardrail that churn among Payroll customers stays at or below 8%. Sequence the attach push so it leans on the error fixes, targeting segments with clean data first.

Integrations: change the unit of success. "Ship 12" measures output, and the number is arbitrary. The 19% of exits citing missing integrations tells us which ones matter. - Objective: Make Quarry the system customers can't easily unplug. - KR1: Share of customers with three or more integrations connected, up from [baseline needed] by an agreed number of points. - KR2: Integration-cited exits down about a third (about 0.9 points of churn). - Ship the integrations ranked by exit and lost-customer demand. That list is the initiative that serves the KRs, not a KR itself.

Growth: the biggest change. Point it at existing customers. New logos and trial conversion don't affect NRR, churn or attach. A 20% rise in new logos can even worsen churn if the new customers are poorly fit. - Objective: Grow revenue from customers we already have. - KR1: NRR from 101% to 108%. Growth owns the expansion motion, and the other teams contribute. - KR2: Price-cited exits down about 15% through packaging and bundling experiments (for example Core HR plus Payroll). Don't raise list prices this quarter, since price is already the second-largest exit reason. - KR3: Share of new customers with three or more integrations connected within 90 days (baseline to be set). This is a lifecycle problem, which is a growth skill. - Trial conversion (9%) stays on the dashboard with a floor, but it is not a KR. If you want new-logo growth kept as a goal, that is a conversation with the CEO, because it contradicts his objective.

Core HR: the weakest link to the company goal. "Best HR platform" can't be scored, and two of the three KRs are ship dates. I have no evidence that the org-chart redesign or custom fields v2 affect retention. They may be worth building, but they don't earn a KR slot without that link. - Objective: Make Core HR the reason customers don't need an all-in-one competitor. - KR1: All-in-one-cited exits down about a third (about 0.5 points of churn). - KR2: Share of customers whose employee data is complete enough to switch on Payroll, up from [baseline needed]. This makes Core HR a direct contributor to attach, and Core HR and Payroll should confirm the definition. - KR3: CSAT from 4.2 to 4.5, kept as a health metric. - Org chart and custom fields stay on the roadmap only if exit or customer data shows they close the competitor gap. Core HR should make that case in week one.

Process recommendations

1. Three objectives per team: decline, with a better answer to "covered." The company has one objective with three KRs. Four teams with three objectives each would create twelve objectives pulling in different directions, which is the problem this memo is fixing. It would also lengthen planning. I recommend one objective and two or three KRs per team. To address the CEO's concern about coverage, give each team a short health-metric list with floors, such as uptime, support SLAs and CSAT. Those get monitored without being goals. A team may request a second objective if it can show a distinct customer outcome that the company objective doesn't cover.

2. OKR scores setting PM bonuses: recommend against. Last year's average score of 0.95 means the goals weren't stretching or the scoring was generous. Either way, the scores carry little information. Tying pay to them would make this worse, because teams would sandbag, avoid ambitious outcomes and game the metrics (ticket counts being the obvious example). The revised KRs are outcome-based and only partly within a team's control, which makes mechanical payouts less fair still.

Instead: - Calibrate so a typical score is around 0.7, with explicit committed and stretch levels. - Use scores as one input to performance reviews, judged by the manager in context. - Revisit a compensation link after two quarters of calibrated scoring.

If the CFO needs something concrete, tie any variable pay to the three company-level KRs, not team scores. That gives a shared incentive to cooperate.

3. Planning took five of thirteen weeks (38% of the quarter): cap it at two. - The CEO fixes company OKRs at least two weeks before the quarter starts, so teams never plan in the dark. - Teams draft on a one-page template with one objective and no more than three KRs, each with a baseline, owner and target. - One 90-minute cross-team session reconciles overlaps, using the exit-reason allocation above as the starting point. - Planning overlaps the last two weeks of the prior quarter and ends with sign-off by day 5 of the new one. - Fewer objectives per team will itself shorten the cycle.

Decisions and next steps

  1. Approve the reframing in this memo, and take the Growth change to the CEO before teams redraft, since it is the one most likely to be contested.
  2. Ask Data for baselines this week: share of customers with three or more integrations, exit reasons by segment, and payroll tickets per pay run. Several KRs above can't be finalised without them.
  3. Ask each team to resubmit within one week.
  4. Take the bonus and three-objective positions to the CEO and CFO together, so we have one conversation rather than two.

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision85%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims52%
  • partialProduces the required deliverable66%
  • passAims the teams with the churn data100%
  • passReads the 0.95 average for what it is99%
  • passKey results are outcomes, not output60%
  • passFocuses on the big rock99%
  • passShows how the goals add up60%
Run
Run
#1
API response time
67 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 1

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic.

GPT-6.1 Sol · API

All statements about the current situation are taken directly from the brief or context, or follow from them; no invented facts.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the supplied context or derived by arithmetic from it.

All got right 10

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

The memo commits to a clear set of recommendations (one objective per team, redirect Growth, keep Payroll alignment) and addresses the process questions unambiguously.

GPT-6.1 Sol · API

The memo commits to clear revised OKRs and process recommendations for Hannah, and notes what would change the call (testing retention effects, validating baselines).

Sonnet 5.5 · API

The memo commits to a clear set of revised team OKRs and process recommendations, and states what would change the call (e.g., CEO conversation on new logos).

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

The output is a memo addressed to Hannah, stays within the 1,200-word limit, and respects the requested form.

GPT-6.1 Sol · API

The output is a memo to Hannah, under 1,200 words, and respects the requested form and reader.

Sonnet 5.5 · API

The output is a memo to Hannah, within the 1,200-word limit, and addresses the requested decisions.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

It names specific unknowns (survey representativeness, systematic differences, missing baselines) and says how to resolve them (test interventions, validate baselines in planning week one).

GPT-6.1 Sol · API

It identifies correlation vs. causation, missing data, and the need to validate targets, and says how to resolve them (retained-revenue bridge, testing, baselines).

Sonnet 5.5 · API

The memo identifies limits in the exit data, selection bias in churn segments, missing baselines, and says how to resolve them in week one.

Avoids unsupported claimsRightRightRight
GPT-6 Astra · ChatGPT

Hypotheses and causal interpretations are clearly labelled as such (e.g., 'priorities, not causal proof', 'we should test'), and no confident claim goes beyond the evidence.

GPT-6.1 Sol · API

Interpretations are labelled as such (correlations, signals, may indicate), and confident claims are supported by the evidence.

Sonnet 5.5 · API

Hypotheses like selection bias and risks are clearly labelled as such; no interpretations are presented as established fact.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

The memo provides revised key results for each team, links them to company goals, and gives actionable process recommendations; a PM could use it with light edits.

GPT-6.1 Sol · API

The memo provides complete revised KRs and process recommendations, within length, and is directly usable by Hannah.

Sonnet 5.5 · API

The memo provides complete revised KRs and process recommendations that Hannah could act on with light edits.

Aims the teams with the churn dataRightRightRight
GPT-6 Astra · ChatGPT

It uses the churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integrations, Core HR at non-Payroll churn, and Growth at price churn, with explicit links.

GPT-6.1 Sol · API

Uses churn reasons and segment data to aim Payroll at errors, Integrations at 3+ integration adoption, and Core HR at churn reduction and price/value risk, with explicit links.

Sonnet 5.5 · API

The memo uses exit survey shares and churn-by-segment data to aim Payroll at errors, Integrations at 3+ integrations, Growth at price, and Core HR at all-in-one competitor.

Reads the 0.95 average for what it isRightRightRight
GPT-6 Astra · ChatGPT

It recognizes the 0.95 average suggests safe targets, and explicitly advises against tying bonuses to OKR scores because it would encourage even safer targets.

GPT-6.1 Sol · API

Recognizes the 0.95 average suggests safe targets, and advises against tying bonuses to OKR scores because it would encourage sandbagging.

Sonnet 5.5 · API

The memo flags the 0.95 average as evidence of safe targets, warns tying bonuses would worsen sandbagging, and recommends against it.

Key results are outcomes, not outputRightRightRight
GPT-6 Astra · ChatGPT

Every revised key result is a measurable change in customer or business behavior (churn rates, attach rates, NRR, CSAT, integration adoption), not a shipping commitment.

GPT-6.1 Sol · API

Every key result is a measurable change in customer or business behavior with a baseline and target; no shipping outputs appear as KRs.

Sonnet 5.5 · API

Every revised key result is a measurable outcome (e.g., error tickets per pay run, exit reductions, attach rate) with baselines and targets; shipping work is explicitly kept as initiatives.

Focuses on the big rockRightRightRight
GPT-6 Astra · ChatGPT

It cuts each team to one objective with a few key results, drops new logos, trial conversion, and shipping goals, and explains why they were dropped.

GPT-6.1 Sol · API

Cuts each team to one objective with a few KRs, drops shipping and acquisition goals, and explains why each was dropped.

Sonnet 5.5 · API

The memo cuts each team to one objective with 2-3 KRs, drops new logos and shipping goals, and explains why they don't serve the company objective.

Shows how the goals add upRightRightRight
GPT-6 Astra · ChatGPT

Each key result is explicitly linked to a company goal (churn, NRR, attach) with reasoning, and the original Growth goals that served no company goal are flagged and replaced.

GPT-6.1 Sol · API

Every KR is linked to a company goal (churn, NRR, attach) with reasoning, and the new-logo goal is explicitly cut because it serves no company goal.

Sonnet 5.5 · API

Every team KR is linked to a company KR (churn, NRR, attach) with arithmetic or reasoning, and the new-logo goal is flagged as serving none and recommended for removal.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT98.5100.03None
3Sonnet 5.5withAPI90.991.72None
4GPT-6 LunawithAPI93.283.32None
5Opus 5.5withClaude90.986.13None
6Gemini 3.8 FlashwithAPI68.270.82None
7Gemini 3.5 Flash-LitewithGemini65.250.03None

About the task

The PM job

Setting a team's goals for the quarter.

Why it matters

Most OKRs are task lists in disguise. Teams ship everything on them and nothing moves.

What good looks like

  • Key results are outcomes, not things to ship
  • Few enough to focus on
  • Each team's goals visibly add up to the company's
  • Kept apart from performance ratings

Deliberately not measured

  • OKR software or formatting
Capability tested

Goal setting

The failure we’re looking for

A task list dressed up as key results

Grading

Decision model and LLM judge, calibrated against a blind PM review