Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 50% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. Makes the decisions clear100% pass
    The final slide gives both decisions with their costs/benefits, an owner, and an 8 October date before Q4 renewals.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  3. A real 7-slide deck100% pass

Where it slips

  1. Owns the problem25% pass
    It owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.
    GPT-6 Astra · ChatGPT · The CEO asks why a metric dropped
  2. Addresses the actual decision54% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Avoids unsupported claims63% pass
    It presents '1,100 customers expect it on 28 Oct' and 'not a saving we planned' as established facts when the source only suggests them.
    Opus 5.5 · Claude · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our CEO, Lena, posted in the leadership channel: “Why did weekly active teams drop 6% last week? Is it the new navigation?” You own the product area. Using the material below, write your reply to Lena, to post in the channel. Keep it under 200 words.

ScenarioTrellis is a project-management tool for small teams. The new navigation went to 100% of web users on Tuesday 15 September after two weeks in a 10% beta. The mobile apps did not change.
MetricsWeekly active teams: week of 7 September 18,420; week of 14 September 17,310 (−6.0%). Web only: −6.3%. Mobile only: −5.2%. Teams that were in the navigation beta: −1.1%. The same week last year, weekly active teams fell 4%, which we put down to the start of the school year.
IncidentOn Thursday 17 September, single sign-on failed for two hours. About 30% of teams use single sign-on.
SupportTickets mentioning navigation: 212 last week, up from 40. Most are some version of “where did X go?”
What a strong answer does

Answers directly that the cause is not established and the navigation is unlikely to be the main driver (mobile, which did not change, fell nearly as much, and the beta cohort barely moved), while taking the navigation support tickets seriously. Names the other candidates with their evidence (seasonality, the SSO outage), says what is being checked and when Lena will hear back.

Critical failures (cap the score)
  • States a single cause as established fact
Case

v1.6 · synthetic · metrics, exec communication

Needs more than a quick edit4 of 10 checks passed

A PM would need to re-check the causal claims and redo the uncertainty handling because the output treats inferences as established and does not specify what would change the call.

Output

Hi Lena,

Good catch—I’ve been digging into the numbers. While the new navigation is a factor, it’s not the whole story.1 Here’s what happened:

  • The SSO Outage: On Thursday, 17 Sept, a two-hour single sign-on failure locked out ~30% of our teams2, directly impacting mid-week engagement.
  • Navigation Friction: Support tickets about navigation jumped from 40 to 212, mostly with “where did X go?” questions. However, teams that were in the earlier beta only dropped 1.1%, suggesting the learning curve is real but manageable once users adjust.3
  • Seasonality: Weekly active teams dropped 4% the exact same week last year due to the start of the school year.4

Next Steps: We’re shipping contextual tooltips for the moved features today and monitoring daily active recovery. I'll share a full breakdown in our standup.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: MetricsMobile only: −5.2%.

    What we’d changeLead with the strongest comparison: mobile fell 5.2% with no navigation change, so the navigation isn't established as a cause. Say what you'll check and when.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    Source: IncidentAbout 30% of teams use single sign-on.

    What we’d change30% of teams use SSO. That doesn't mean all of them were trying to get in during those two hours.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeBeta teams may differ from everyone else. Their small drop doesn't show a manageable learning curve.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    Source: Metricswhich we put down to the start of the school year.

    What we’d changeLast year's cause was an attribution, not a finding. Say 'which we put down to'.

Check by check

Got wrong · 6
  • Uses the supplied evidence correctlyIt introduces causal statements such as 'directly impacting mid-week engagement' and 'new navigation is a factor' that are not in the supplied context or derivable by arithmetic.
  • Addresses the actual decisionIt does not say what result or condition would change the answer; it only says it will monitor and share a breakdown.
  • Identifies material uncertaintyIt does not name the specific unknowns that could change the decision or tie the monitoring to what would settle the call.
  • Avoids unsupported claimsIt presents causal interpretations such as the SSO outage directly impacting engagement and navigation being a factor as established fact, though the evidence only suggests them.
  • Separates effect from explanationIt lists SSO outage, navigation friction, and seasonality as 'what happened' rather than clearly labelling them as hypotheses with supporting or contradicting evidence.
  • Owns the problemIt reports the SSO outage but does not say what the team has already done about it or recommend a way forward for that specific problem.
Got right · 4
  • Respects explicit constraintsIt is a channel reply addressed to Lena and is under 200 words.
  • Produces the required deliverableThe reply is in the right channel format, addressed to Lena, under 200 words, and usable as a status update.
  • Answers the question askedThe first lines answer the question by saying the new navigation is a factor but not the whole story.
  • Ready for the reader as deliveredIt is written for Lena, uses no placeholders, and all dates and owners are taken from the material or clearly proposed.

Claims the judge couldn’t find in the brief

  • The new navigation is a factor, though not the whole story.
  • The SSO outage directly impacted mid-week engagement.
  • This suggests the learning curve is real but manageable once users adjust.

Grades and run details

Decision model 55 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly38%
  • partialAddresses the actual decision73%
  • passRespects explicit constraints55%
  • partialIdentifies material uncertainty69%
  • failAvoids unsupported claims23%
  • passProduces the required deliverable57%
  • partialAnswers the question asked32%
  • partialSeparates effect from explanation68%
  • partialOwns the problem94%
  • passReady for the reader as delivered84%
Run
Run
#1
Time to output
5 s
Submitted
27 Sept 2026

Usable with a quick edit9 of 10 checks passed

The reasoning and evidence are sound; only a minor addition of already-taken actions or a trade-off is needed.

Output

Lena — we don’t yet have enough evidence to attribute the 6% drop to navigation. My read is that it may be contributing, but it’s unlikely to be the whole story:

  • Mobile-only teams fell 5.2% despite no navigation change, versus 6.3% for web-only teams.
  • Navigation-beta teams fell just 1.1%. That’s reassuring, but they may not be representative.
  • The same week last year fell 4%, suggesting a seasonal component—not proving one.
  • Thursday’s two-hour SSO outage is another plausible contributor; about 30% of teams use SSO.

There is clear navigation friction: tickets rose from 40 to 212, mostly “where did X go?” That warrants action even if it doesn’t explain the activity decline.

I’ll break down daily activity by platform, SSO usage, and beta exposure, and check whether key navigation-dependent workflows deteriorated after Tuesday’s rollout. We’ll also prioritize fixes for the most common findability complaints.

I’ll share an initial read tomorrow, including whether the evidence supports a rollback. I wouldn’t treat seasonality, the outage, and navigation as additive explanations without that analysis.

Check by check

Got wrong · 1
  • Owns the problemIt proposes future analysis and fixes but does not state any actions the team has already taken or offer trade-offs.
Got right · 9
  • Uses the supplied evidence correctlyAll factual numbers, dates, and quotes are drawn accurately from the supplied context, with interpretations clearly hedged.
  • Addresses the actual decisionThe output commits early that navigation is unlikely to be the whole story and says tomorrow's analysis will show whether rollback is supported.
  • Respects explicit constraintsThe reply is under 200 words, addressed to Lena, and fits a leadership-channel update.
  • Identifies material uncertaintyIt names the non-representative beta cohort, unproven seasonality, and SSO outage as unknowns and outlines the analysis that will resolve them.
  • Avoids unsupported claimsCauses are labelled as possible ('may be contributing', 'plausible', 'suggesting') rather than asserted as fact.
  • Produces the required deliverableIt is a complete channel reply with an answer, evidence, and next steps, within the requested length.
  • Answers the question askedThe first line directly addresses whether the new navigation caused the drop.
  • Separates effect from explanationEach candidate explanation is framed as a hypothesis with supporting or contradicting evidence (mobile unchanged, beta low, seasonality, SSO).
  • Ready for the reader as deliveredIt is written entirely for Lena with no placeholders or unexplained internal labels.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly74%
  • passAddresses the actual decision31%
  • passRespects explicit constraints61%
  • partialIdentifies material uncertainty70%
  • passAvoids unsupported claims79%
  • passProduces the required deliverable97%
  • passAnswers the question asked77%
  • passSeparates effect from explanation98%
  • partialOwns the problem83%
  • passReady for the reader as delivered97%
Run
Run
#1
API response time
8 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT87.387.12None
2GPT-6 LunawithAPI81.082.62None
3GPT-6.1 SolwithAPI77.783.02None
4Sonnet 5.5withAPI78.564.42None
5Opus 5.5withClaude80.451.92None
6Gemini 3.5 Flash-LitewithGemini62.139.02None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours