Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 43% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. A real 7-slide deck100% pass
  3. Answers the question asked96% pass
    The first line directly addresses whether the new navigation caused the drop.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped

Where it slips

  1. Owns the problem23% pass
    It proposes future analysis and fixes but does not state any actions the team has already taken or offer trade-offs.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped
  2. Addresses the actual decision41% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Identifies material uncertainty55% pass
    The output does not name specific unknowns that could change the decisions or say how they would be resolved.
    GPT-6 Luna · API · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our CEO, Lena, posted in the leadership channel: “Why did weekly active teams drop 6% last week? Is it the new navigation?” You own the product area. Using the material below, write your reply to Lena, to post in the channel. Keep it under 200 words.

What the model was given4 items: Scenario, Metrics, Incident, Support
ScenarioTrellis is a project-management tool for small teams. The new navigation went to 100% of web users on Tuesday 15 September after two weeks in a 10% beta. The mobile apps did not change.
MetricsWeekly active teams: week of 7 September 18,420; week of 14 September 17,310 (−6.0%). Web only: −6.3%. Mobile only: −5.2%. Teams that were in the navigation beta: −1.1%. The same week last year, weekly active teams fell 4%, which we put down to the start of the school year.
IncidentOn Thursday 17 September, single sign-on failed for two hours. About 30% of teams use single sign-on.
SupportTickets mentioning navigation: 212 last week, up from 40. Most are some version of “where did X go?”
What a strong answer doesThe answer key the graders mark against

Answers directly that the cause is not established and the navigation is unlikely to be the main driver (mobile, which did not change, fell nearly as much, and the beta cohort barely moved), while taking the navigation support tickets seriously. Names the other candidates with their evidence (seasonality, the SSO outage), says what is being checked and when Lena will hear back.

Critical failures (cap the score at 40)
  • States a single cause as established fact
Case

Mid-level · v1.8 · synthetic · metrics, exec communication

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit1 of 10 checks passed

A PM would need to re-check evidence and rewrite the causal claims, uncertainty, and next steps. The invented tooltip deployment and overconfident navigation conclusion are not quick wording fixes.

Output

Hi Lena — the new navigation isn't the primary driver of the 6% drop.

The data points to broader seasonal and technical factors:

  • Seasonality: We historically see a dip this week (-4% last year) due to the start of the school year.
  • Mobile dropped similarly: Mobile-only activity fell 5.2% despite receiving zero navigation changes (vs. -6.3% on web).
  • Beta teams held steady: Teams familiar with the new UI only dropped 1.1%.
  • SSO outage: Thursday’s 2-hour outage locked out access for up to 30% of our user base, contributing to the dip.

There is some expected initial friction: navigation support tickets rose from 40 to 212, primarily around discoverability (“where did X go?”).

Next steps: We are deploying targeted in-app tooltips this week to guide users to relocated features and expect activity to rebound as users adjust and seasonal trends normalize. I’ll share an updated cohort analysis on Friday.

Claims the judge couldn’t find in the brief

  • The new navigation isn't the primary driver of the 6% drop.
  • The data points to broader seasonal and technical factors.
  • We historically see a dip this week (-4% last year) due to the start of the school year.
  • Thursday’s 2-hour outage locked out access for up to 30% of our user base.
  • The SSO outage contributed to the dip.
  • There is some expected initial friction.
  • We are deploying targeted in-app tooltips this week.

Grades and run details

Decision model 55 · LLM judge 1 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly34%
  • partialAddresses the actual decision74%
  • passRespects explicit constraints77%
  • failIdentifies material uncertainty36%
  • failAvoids unsupported claims25%
  • passProduces the required deliverable96%
  • passAnswers the question asked99%
  • partialSeparates effect from explanation32%
  • partialOwns the problem87%
  • passReady for the reader as delivered50%
Run
Run
#1
API response time
17 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 10 checks passed

A PM would need to verify the weekly-active metric definition and add explicit conditions that would change the conclusion before posting; that is more than a quick edit.

Output

Why I don't think the nav is the main cause

Mobile didn't change and still fell 5.2%. Web fell 6.3%. So about 5 points hit everyone, and the nav can explain at most about 1 point on web.1 Seasonality likely covers most of the shared drop. The same week last year we fell 4% as the school year started. The SSO outage (Thursday, 2 hours, about 30% of teams) is probably minor. A team only needs one session in the week to count as active.2 I'll confirm by comparing SSO and non-SSO teams. Beta teams fell only 1.1%. That looks reassuring, but beta teams may be more engaged than average, so I'm not relying on it.

Where the nav is hurting Navigation tickets rose from 40 to 212, and most are "where did X go?" That's friction now and a churn risk if we leave it.

Next steps, by Wednesday

Split the drop by SSO vs non-SSO teams, and beta vs comparable non-beta teams. Ship quick fixes for the most-asked "where did X go" items. Watch this week's number. If the drop is seasonal, it should flatten.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: MetricsWeb only: −6.3%. Mobile only: −5.2%.

    What we’d changeDon't cap the navigation's effect at 1 point: web and mobile differ in more than the navigation, so the subtraction doesn't hold. Give seasonality less credit, and hold the SSO reasoning until the split is done.

  2. 2Invented evidenceVerify or remove the claimQuick edit

    What we’d changeThe brief doesn't define an active team. Present it as an assumption to check, since the SSO argument depends on it.

Claims the judge couldn’t find in the brief

  • A team only needs one session in the week to count as active.

Grades and run details

Decision model 80 · LLM judge 5 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly26%
  • passAddresses the actual decision34%
  • passRespects explicit constraints70%
  • passIdentifies material uncertainty91%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable95%
  • passAnswers the question asked98%
  • passSeparates effect from explanation85%
  • partialOwns the problem80%
  • passReady for the reader as delivered68%
Run
Run
#1
Time to output
30 s
Submitted
27 Sept 2026

Needs more than a quick edit5 of 10 checks passed

It includes an unsupported SSO recovery assumption, so evidence must be re-checked. It also lacks an explicit condition that would change the call, so a PM would need to add that.

Output

Lena, short answer: partly, but navigation probably isn't the main driver.

  • Seasonality is likely most of it. Same week last year we fell 4% with the school-year restart. That leaves about 2 points unexplained.
  • Mobile fell 5.2% and didn't change. Web fell 6.3%, so the navigation's effect is roughly 1 point at most.
  • Beta teams fell only 1.1%. That supports a small navigation effect, but the beta group may be more engaged than average, so I wouldn't lean on it.
  • Support tickets mentioning navigation rose from 40 to 212, mostly "where did X go?" That's real friction, but small next to about 1,100 lost teams.
  • The SSO outage on Thursday lasted two hours and affected about 30% of teams. Most would have come back later in the week, so I doubt it explains much.

I can't cleanly separate these yet. This week I'll: 1. Break the drop down by SSO vs. non-SSO teams and by new vs. established teams. 2. Check whether last week's inactive teams return this week. 3. Fix the top "where did X go?" items.

I'll post an update Friday.

Claims the judge couldn’t find in the brief

  • Most SSO-affected teams would have come back later in the week.
  • Navigation support-ticket friction is small next to about 1,100 lost teams.

Grades and run details

Decision model 80 · LLM judge 5 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly12%
  • partialAddresses the actual decision52%
  • passRespects explicit constraints63%
  • passIdentifies material uncertainty47%
  • failAvoids unsupported claims30%
  • passProduces the required deliverable96%
  • passAnswers the question asked70%
  • passSeparates effect from explanation82%
  • partialOwns the problem88%
  • passReady for the reader as delivered82%
Run
Run
#1
API response time
9 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongWrongMixed
Gemini 3.8 Flash · API

It invents or overstates several current-situation facts, including user-base lockout, causal contribution, historical seasonality, and an ongoing tooltip deployment.

Opus 5.5 · Claude

It introduces an unsupported definition that a weekly active team only needs one session, which is not in the supplied context and does not follow by arithmetic.

Sonnet 5.5 · API

The SSO recovery claim ('Most would have come back later in the week') is not in the brief or derivable from it, so not all current-situation statements are supported.

Addresses the actual decisionWrongMixedWrong
Gemini 3.8 Flash · API

It commits early that navigation is not the primary driver but does not say what result or condition would change that answer.

Opus 5.5 · Claude

It commits early that the nav is not the main cause but never states what result or condition would overturn that call.

Sonnet 5.5 · API

It commits to 'partly, but navigation probably isn't the main driver' but never states what result or condition would change that call, only lists analyses.

Respects explicit constraintsMixedRightRight
Gemini 3.8 Flash · API

Although it is addressed to Lena and under 200 words, it violates the supplied-material constraint by adding unsupported current actions and causal claims.

Opus 5.5 · Claude

The reply is under 200 words and addressed as a channel reply to Lena.

Sonnet 5.5 · API

At about 180 words and addressed to Lena in channel-reply form, it respects the length and reader constraints.

Identifies material uncertaintyWrongMixedMixed
Gemini 3.8 Flash · API

It does not name the material unknowns or explain how the Friday cohort analysis would change the call.

Opus 5.5 · Claude

It names unknowns and says how to investigate but does not specify what outcomes would change the decision.

Sonnet 5.5 · API

It names analyses to run but does not identify the specific results that would change the call, and the SSO impact is asserted rather than treated as unknown.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

The reply is not usable as-is because it overstates causality, omits the required uncertainty framing, and includes invented actions.

Opus 5.5 · Claude

The deliverable is a complete, usable channel reply within length.

Sonnet 5.5 · API

The reply is complete and usable as a channel update, with an answer, evidence, and next steps.

Separates effect from explanationWrongRightRight
Gemini 3.8 Flash · API

Possible causes are presented as data-backed factors rather than labelled hypotheses with evidence and contradictions.

Opus 5.5 · Claude

Causes are labelled with uncertainty and paired with supporting or contradicting evidence.

Sonnet 5.5 · API

Seasonality, navigation, support friction, and SSO are each labelled as hypotheses with supporting or contradicting evidence.

Ready for the reader as deliveredMixedRightRight
Gemini 3.8 Flash · API

It includes invented current actions and unsupported causal statements, so it should not go to Lena without substantive correction.

Opus 5.5 · Claude

Written for Lena with no unclear placeholders or lifted internal wording.

Sonnet 5.5 · API

Written for Lena, no placeholders, and names and dates are from the material or clearly proposed.

All got wrong 2

Avoids unsupported claimsWrongWrongWrong
Gemini 3.8 Flash · API

It presents navigation not being primary, SSO contributing, and expected rebound as established or near-established rather than labelled hypotheses.

Opus 5.5 · Claude

It presents the one-session active-definition claim as an established fact with no support.

Sonnet 5.5 · API

It presents 'Most would have come back later in the week' and 'small next to about 1,100 lost teams' as fact without support.

Owns the problemWrongWrongWrong
Gemini 3.8 Flash · API

It proposes tooltips and an analysis but does not provide already-taken actions with a recommended option and trade-off for each material problem.

Opus 5.5 · Claude

It reports problems but does not state what the team has already done about them before recommending next steps.

Sonnet 5.5 · API

It reports support-ticket friction and SSO outage but does not say what the team has already done, only future fixes and analyses.

All got right 1

Answers the question askedRightRightRight
Gemini 3.8 Flash · API

The first line directly answers whether the new navigation caused the drop.

Opus 5.5 · Claude

The first lines answer that the navigation is not the main cause.

Sonnet 5.5 · API

It answers in the first line that navigation is partly responsible but probably not the main driver.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 80% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT86.787.12None
2GPT-6.1 SolwithAPI75.883.02None
3GPT-6 LunawithAPI73.882.62None
4Opus 5.5withClaude80.451.92None
5Sonnet 5.5withAPI68.864.42None
6Gemini 3.5 Flash-LitewithGemini62.139.02None
7Gemini 3.8 FlashwithAPI44.837.92None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours