Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 43% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. A real 7-slide deck100% pass
  3. Answers the question asked96% pass
    The first line directly addresses whether the new navigation caused the drop.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped

Where it slips

  1. Owns the problem23% pass
    It proposes future analysis and fixes but does not state any actions the team has already taken or offer trade-offs.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped
  2. Addresses the actual decision41% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Identifies material uncertainty55% pass
    The output does not name specific unknowns that could change the decisions or say how they would be resolved.
    GPT-6 Luna · API · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our CEO, Lena, posted in the leadership channel: “Why did weekly active teams drop 6% last week? Is it the new navigation?” You own the product area. Using the material below, write your reply to Lena, to post in the channel. Keep it under 200 words.

What the model was given4 items: Scenario, Metrics, Incident, Support
ScenarioTrellis is a project-management tool for small teams. The new navigation went to 100% of web users on Tuesday 15 September after two weeks in a 10% beta. The mobile apps did not change.
MetricsWeekly active teams: week of 7 September 18,420; week of 14 September 17,310 (−6.0%). Web only: −6.3%. Mobile only: −5.2%. Teams that were in the navigation beta: −1.1%. The same week last year, weekly active teams fell 4%, which we put down to the start of the school year.
IncidentOn Thursday 17 September, single sign-on failed for two hours. About 30% of teams use single sign-on.
SupportTickets mentioning navigation: 212 last week, up from 40. Most are some version of “where did X go?”
What a strong answer doesThe answer key the graders mark against

Answers directly that the cause is not established and the navigation is unlikely to be the main driver (mobile, which did not change, fell nearly as much, and the beta cohort barely moved), while taking the navigation support tickets seriously. Names the other candidates with their evidence (seasonality, the SSO outage), says what is being checked and when Lena will hear back.

Critical failures (cap the score at 40)
  • States a single cause as established fact
Case

Mid-level · v1.8 · synthetic · metrics, exec communication

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 10 checks passed

A PM would need to re-check the causal claims and redo the uncertainty handling because the output treats inferences as established and does not specify what would change the call.

Output

Hi Lena,

Good catch—I’ve been digging into the numbers. While the new navigation is a factor, it’s not the whole story.1 Here’s what happened:

  • The SSO Outage: On Thursday, 17 Sept, a two-hour single sign-on failure locked out ~30% of our teams2, directly impacting mid-week engagement.
  • Navigation Friction: Support tickets about navigation jumped from 40 to 212, mostly with “where did X go?” questions. However, teams that were in the earlier beta only dropped 1.1%, suggesting the learning curve is real but manageable once users adjust.3
  • Seasonality: Weekly active teams dropped 4% the exact same week last year due to the start of the school year.4

Next Steps: We’re shipping contextual tooltips for the moved features today and monitoring daily active recovery. I'll share a full breakdown in our standup.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: MetricsMobile only: −5.2%.

    What we’d changeLead with the strongest comparison: mobile fell 5.2% with no navigation change, so the navigation isn't established as a cause. Say what you'll check and when.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    Source: IncidentAbout 30% of teams use single sign-on.

    What we’d change30% of teams use SSO. That doesn't mean all of them were trying to get in during those two hours.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeBeta teams may differ from everyone else. Their small drop doesn't show a manageable learning curve.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    Source: Metricswhich we put down to the start of the school year.

    What we’d changeLast year's cause was an attribution, not a finding. Say 'which we put down to'.

Claims the judge couldn’t find in the brief

  • The new navigation is a factor, though not the whole story.
  • The SSO outage directly impacted mid-week engagement.
  • This suggests the learning curve is real but manageable once users adjust.

Grades and run details

Decision model 55 · LLM judge 4 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly51%
  • partialAddresses the actual decision40%
  • passRespects explicit constraints49%
  • partialIdentifies material uncertainty58%
  • failAvoids unsupported claims32%
  • passProduces the required deliverable50%
  • partialAnswers the question asked37%
  • partialSeparates effect from explanation63%
  • partialOwns the problem95%
  • passReady for the reader as delivered48%
Run
Run
#1
Time to output
5 s
Submitted
27 Sept 2026

Usable with a quick edit8 of 10 checks passed

The evidence and reasoning are correct; only small omissions like a timeline and already-taken actions would need a quick edit.

Output

Lena, we can’t yet attribute the 6% drop to the new navigation. The timing and jump in navigation-related tickets (212, up from 40) make discoverability a real concern, but teams in the two-week beta were down only 1.1%. The decline also appears on mobile (−5.2%), which didn’t change, so navigation alone is unlikely to explain it.

Two other factors may contribute: WAU fell 4% in the same week last year, around the start of the school year, and SSO was unavailable for two hours on Thursday for teams using it (about 30%).

I’d treat the navigation as a usability issue to address, not yet the proven cause of the overall decline. We’ll compare affected and unaffected cohorts, check the SSO impact, and investigate where users are getting stuck; meanwhile, we’ll prioritize fixes for the “where did X go?” feedback.

Grades and run details

Decision model 90 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly69%
  • partialAddresses the actual decision29%
  • passRespects explicit constraints32%
  • passIdentifies material uncertainty62%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable53%
  • passAnswers the question asked90%
  • passSeparates effect from explanation84%
  • partialOwns the problem92%
  • passReady for the reader as delivered91%
Run
Run
#1
API response time
6 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 10 checks passed

The evidence and reasoning are sound and the reply is post-ready. A quick edit could add already-taken actions or trade-offs, but no evidence re-check or rework is needed.

Output

Lena, we lost 1,110 weekly active teams. Navigation may be contributing, but we don’t yet know how much.1 I own investigating this and the response.

Mobile-only teams also fell 5.2%, despite no navigation change, so there’s likely a broader factor. Last year’s 4% dip makes seasonality plausible, but doesn’t establish this year’s cause. Thursday’s two-hour SSO outage is another candidate; 30% of teams use SSO, but that doesn’t tell us how many missed the entire week.

Navigation tickets rose from 40 to 212, mostly “where did X go?” That’s clear friction we should address now, though it doesn’t establish the cause of the activity drop. Beta teams fell just 1.1%; familiarity could help, but they may differ from other teams.

I’ll report back tomorrow with daily activity around Tuesday’s rollout and Thursday’s outage, split by platform, SSO use and beta status, checking tracking and prior trends. I’ll also check whether affected teams returned later in the week and where key workflows deteriorated.

Meanwhile, we’ll prioritise the missing destinations behind those tickets. I wouldn’t roll back solely on this aggregate result; evidence of navigation blocking core work would change that.

What a PM had to fix

  1. 1Decision deferredMake the callQuick edit

    Source: MetricsMobile only: −5.2%.

    What we’d changeAnswer Lena's question more directly: mobile's 5.2% drop with no navigation change points away from the navigation as the main cause. Tighten the SSO hedge.

Grades and run details

Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly41%
  • failAddresses the actual decision19%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims55%
  • passProduces the required deliverable95%
  • passAnswers the question asked41%
  • passSeparates effect from explanation98%
  • partialOwns the problem71%
  • passReady for the reader as delivered86%
Run
Run
#1
Time to output
26 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It introduces causal statements such as 'directly impacting mid-week engagement' and 'new navigation is a factor' that are not in the supplied context or derivable by arithmetic.

GPT-6 Luna · API

All named facts, figures, and quotes match the supplied context with no invented current-state information.

GPT-6 Astra · ChatGPT

All stated facts and figures come directly from the supplied context or follow by arithmetic, with no invented current-situation claims.

Addresses the actual decisionWrongWrongMixed
Gemini 3.5 Flash-Lite · Gemini

It does not say what result or condition would change the answer; it only says it will monitor and share a breakdown.

GPT-6 Luna · API

It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.

GPT-6 Astra · ChatGPT

It commits clearly to 'navigation may be contributing but we don't yet know how much' and says evidence of navigation blocking core work would change the rollback decision.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It does not name the specific unknowns that could change the decision or tie the monitoring to what would settle the call.

GPT-6 Luna · API

It names the unknown cause, bounds navigation with beta and mobile data, and states it will compare cohorts and check SSO impact.

GPT-6 Astra · ChatGPT

It names the key unknowns—navigation contribution, broader factor, seasonality, SSO impact, beta selection—and says daily activity split by platform, SSO and beta will resolve them.

Avoids unsupported claimsWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It presents causal interpretations such as the SSO outage directly impacting engagement and navigation being a factor as established fact, though the evidence only suggests them.

GPT-6 Luna · API

Causal statements are hedged as concerns or possibilities and are tied to supplied evidence.

GPT-6 Astra · ChatGPT

Hypotheses are labelled with 'may,' 'likely,' 'plausible,' 'candidate,' and 'could,' and no unlabelled cause is asserted as fact.

Separates effect from explanationWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It lists SSO outage, navigation friction, and seasonality as 'what happened' rather than clearly labelling them as hypotheses with supporting or contradicting evidence.

GPT-6 Luna · API

Navigation, seasonality, and SSO are treated as possible explanations, each with supporting or contradicting evidence.

GPT-6 Astra · ChatGPT

Navigation, seasonality, SSO outage, and beta learnings are each treated as possible explanations with their supporting or contradicting evidence.

All got wrong 1

Owns the problemWrongWrongWrong
Gemini 3.5 Flash-Lite · Gemini

It reports the SSO outage but does not say what the team has already done about it or recommend a way forward for that specific problem.

GPT-6 Luna · API

It recommends next steps but does not state what the team has already done about the reported problems or offer trade-offs.

GPT-6 Astra · ChatGPT

It owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.

All got right 4

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It is a channel reply addressed to Lena and is under 200 words.

GPT-6 Luna · API

It is a channel reply to Lena, under 200 words, and uses the supplied material.

GPT-6 Astra · ChatGPT

It is addressed to Lena, written as a channel reply, and is 189 words, under the 200-word limit.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The reply is in the right channel format, addressed to Lena, under 200 words, and usable as a status update.

GPT-6 Luna · API

The required reply is present, complete, under 200 words, and usable as a leadership-channel update.

GPT-6 Astra · ChatGPT

The requested reply is present and complete, under the word limit, and usable by Lena as a channel update.

Answers the question askedRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The first lines answer the question by saying the new navigation is a factor but not the whole story.

GPT-6 Luna · API

The first line directly answers that the drop cannot yet be attributed to the new navigation.

GPT-6 Astra · ChatGPT

The first lines directly answer whether the new navigation caused the drop by saying it may be contributing but the amount is not yet known.

Ready for the reader as deliveredRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It is written for Lena, uses no placeholders, and all dates and owners are taken from the material or clearly proposed.

GPT-6 Luna · API

It is written entirely for Lena with no placeholders, unexplained terms, or invented owners or dates.

GPT-6 Astra · ChatGPT

It is written entirely for Lena with no unexplained brief terms, placeholders, or invented owners or dates.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 80% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT86.787.12None
2GPT-6.1 SolwithAPI75.883.02None
3GPT-6 LunawithAPI73.882.62None
4Opus 5.5withClaude80.451.92None
5Sonnet 5.5withAPI68.864.42None
6Gemini 3.5 Flash-LitewithGemini62.139.02None
7Gemini 3.8 FlashwithAPI44.837.92None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours