Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 50% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. Makes the decisions clear100% pass
    The final slide gives both decisions with their costs/benefits, an owner, and an 8 October date before Q4 renewals.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  3. A real 7-slide deck100% pass

Where it slips

  1. Owns the problem25% pass
    It owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.
    GPT-6 Astra · ChatGPT · The CEO asks why a metric dropped
  2. Addresses the actual decision54% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Avoids unsupported claims63% pass
    It presents '1,100 customers expect it on 28 Oct' and 'not a saving we planned' as established facts when the source only suggests them.
    Opus 5.5 · Claude · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our CEO, Lena, posted in the leadership channel: “Why did weekly active teams drop 6% last week? Is it the new navigation?” You own the product area. Using the material below, write your reply to Lena, to post in the channel. Keep it under 200 words.

ScenarioTrellis is a project-management tool for small teams. The new navigation went to 100% of web users on Tuesday 15 September after two weeks in a 10% beta. The mobile apps did not change.
MetricsWeekly active teams: week of 7 September 18,420; week of 14 September 17,310 (−6.0%). Web only: −6.3%. Mobile only: −5.2%. Teams that were in the navigation beta: −1.1%. The same week last year, weekly active teams fell 4%, which we put down to the start of the school year.
IncidentOn Thursday 17 September, single sign-on failed for two hours. About 30% of teams use single sign-on.
SupportTickets mentioning navigation: 212 last week, up from 40. Most are some version of “where did X go?”
What a strong answer does

Answers directly that the cause is not established and the navigation is unlikely to be the main driver (mobile, which did not change, fell nearly as much, and the beta cohort barely moved), while taking the navigation support tickets seriously. Names the other candidates with their evidence (seasonality, the SSO outage), says what is being checked and when Lena will hear back.

Critical failures (cap the score)
  • States a single cause as established fact
Case

v1.6 · synthetic · metrics, exec communication

Usable with a quick edit8 of 10 checks passed

The evidence and reasoning are correct; only small omissions like a timeline and already-taken actions would need a quick edit.

Output

Lena, we can’t yet attribute the 6% drop to the new navigation. The timing and jump in navigation-related tickets (212, up from 40) make discoverability a real concern, but teams in the two-week beta were down only 1.1%. The decline also appears on mobile (−5.2%), which didn’t change, so navigation alone is unlikely to explain it.

Two other factors may contribute: WAU fell 4% in the same week last year, around the start of the school year, and SSO was unavailable for two hours on Thursday for teams using it (about 30%).

I’d treat the navigation as a usability issue to address, not yet the proven cause of the overall decline. We’ll compare affected and unaffected cohorts, check the SSO impact, and investigate where users are getting stuck; meanwhile, we’ll prioritize fixes for the “where did X go?” feedback.

Check by check

Got wrong · 2
  • Addresses the actual decisionIt commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
  • Owns the problemIt recommends next steps but does not state what the team has already done about the reported problems or offer trade-offs.
Got right · 8
  • Uses the supplied evidence correctlyAll named facts, figures, and quotes match the supplied context with no invented current-state information.
  • Respects explicit constraintsIt is a channel reply to Lena, under 200 words, and uses the supplied material.
  • Identifies material uncertaintyIt names the unknown cause, bounds navigation with beta and mobile data, and states it will compare cohorts and check SSO impact.
  • Avoids unsupported claimsCausal statements are hedged as concerns or possibilities and are tied to supplied evidence.
  • Produces the required deliverableThe required reply is present, complete, under 200 words, and usable as a leadership-channel update.
  • Answers the question askedThe first line directly answers that the drop cannot yet be attributed to the new navigation.
  • Separates effect from explanationNavigation, seasonality, and SSO are treated as possible explanations, each with supporting or contradicting evidence.
  • Ready for the reader as deliveredIt is written entirely for Lena with no placeholders, unexplained terms, or invented owners or dates.

Grades and run details

Decision model 85 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly75%
  • partialAddresses the actual decision62%
  • passRespects explicit constraints32%
  • partialIdentifies material uncertainty90%
  • passAvoids unsupported claims77%
  • passProduces the required deliverable45%
  • passAnswers the question asked87%
  • passSeparates effect from explanation84%
  • partialOwns the problem92%
  • passReady for the reader as delivered95%
Run
Run
#1
API response time
6 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT87.387.12None
2GPT-6 LunawithAPI81.082.62None
3GPT-6.1 SolwithAPI77.783.02None
4Sonnet 5.5withAPI78.564.42None
5Opus 5.5withClaude80.451.92None
6Gemini 3.5 Flash-LitewithGemini62.139.02None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours