Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 50% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. Makes the decisions clear100% pass
    The final slide gives both decisions with their costs/benefits, an owner, and an 8 October date before Q4 renewals.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  3. A real 7-slide deck100% pass

Where it slips

  1. Owns the problem25% pass
    It owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.
    GPT-6 Astra · ChatGPT · The CEO asks why a metric dropped
  2. Addresses the actual decision54% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Avoids unsupported claims63% pass
    It presents '1,100 customers expect it on 28 Oct' and 'not a saving we planned' as established facts when the source only suggests them.
    Opus 5.5 · Claude · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our CEO, Lena, posted in the leadership channel: “Why did weekly active teams drop 6% last week? Is it the new navigation?” You own the product area. Using the material below, write your reply to Lena, to post in the channel. Keep it under 200 words.

ScenarioTrellis is a project-management tool for small teams. The new navigation went to 100% of web users on Tuesday 15 September after two weeks in a 10% beta. The mobile apps did not change.
MetricsWeekly active teams: week of 7 September 18,420; week of 14 September 17,310 (−6.0%). Web only: −6.3%. Mobile only: −5.2%. Teams that were in the navigation beta: −1.1%. The same week last year, weekly active teams fell 4%, which we put down to the start of the school year.
IncidentOn Thursday 17 September, single sign-on failed for two hours. About 30% of teams use single sign-on.
SupportTickets mentioning navigation: 212 last week, up from 40. Most are some version of “where did X go?”
What a strong answer does

Answers directly that the cause is not established and the navigation is unlikely to be the main driver (mobile, which did not change, fell nearly as much, and the beta cohort barely moved), while taking the navigation support tickets seriously. Names the other candidates with their evidence (seasonality, the SSO outage), says what is being checked and when Lena will hear back.

Critical failures (cap the score)
  • States a single cause as established fact
Case

v1.6 · synthetic · metrics, exec communication

Usable with a quick edit9 of 10 checks passed

The evidence and reasoning are sound and the reply is post-ready. A quick edit could add already-taken actions or trade-offs, but no evidence re-check or rework is needed.

Output

Lena, we lost 1,110 weekly active teams. Navigation may be contributing, but we don’t yet know how much.1 I own investigating this and the response.

Mobile-only teams also fell 5.2%, despite no navigation change, so there’s likely a broader factor. Last year’s 4% dip makes seasonality plausible, but doesn’t establish this year’s cause. Thursday’s two-hour SSO outage is another candidate; 30% of teams use SSO, but that doesn’t tell us how many missed the entire week.

Navigation tickets rose from 40 to 212, mostly “where did X go?” That’s clear friction we should address now, though it doesn’t establish the cause of the activity drop. Beta teams fell just 1.1%; familiarity could help, but they may differ from other teams.

I’ll report back tomorrow with daily activity around Tuesday’s rollout and Thursday’s outage, split by platform, SSO use and beta status, checking tracking and prior trends. I’ll also check whether affected teams returned later in the week and where key workflows deteriorated.

Meanwhile, we’ll prioritise the missing destinations behind those tickets. I wouldn’t roll back solely on this aggregate result; evidence of navigation blocking core work would change that.

What a PM had to fix

  1. 1Decision deferredMake the callQuick edit

    Source: MetricsMobile only: −5.2%.

    What we’d changeAnswer Lena's question more directly: mobile's 5.2% drop with no navigation change points away from the navigation as the main cause. Tighten the SSO hedge.

Check by check

Got wrong · 1
  • Owns the problemIt owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.
Got right · 9
  • Uses the supplied evidence correctlyAll stated facts and figures come directly from the supplied context or follow by arithmetic, with no invented current-situation claims.
  • Addresses the actual decisionIt commits clearly to 'navigation may be contributing but we don't yet know how much' and says evidence of navigation blocking core work would change the rollback decision.
  • Respects explicit constraintsIt is addressed to Lena, written as a channel reply, and is 189 words, under the 200-word limit.
  • Identifies material uncertaintyIt names the key unknowns—navigation contribution, broader factor, seasonality, SSO impact, beta selection—and says daily activity split by platform, SSO and beta will resolve them.
  • Avoids unsupported claimsHypotheses are labelled with 'may,' 'likely,' 'plausible,' 'candidate,' and 'could,' and no unlabelled cause is asserted as fact.
  • Produces the required deliverableThe requested reply is present and complete, under the word limit, and usable by Lena as a channel update.
  • Answers the question askedThe first lines directly answer whether the new navigation caused the drop by saying it may be contributing but the amount is not yet known.
  • Separates effect from explanationNavigation, seasonality, SSO outage, and beta learnings are each treated as possible explanations with their supporting or contradicting evidence.
  • Ready for the reader as deliveredIt is written entirely for Lena with no unexplained brief terms, placeholders, or invented owners or dates.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly52%
  • partialAddresses the actual decision46%
  • passRespects explicit constraints43%
  • passIdentifies material uncertainty59%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable95%
  • passAnswers the question asked49%
  • passSeparates effect from explanation98%
  • partialOwns the problem69%
  • passReady for the reader as delivered96%
Run
Run
#1
Time to output
26 s
Submitted
25 Sept 2026

Needs more than a quick edit5 of 10 checks passed

It includes an unsupported SSO recovery assumption, so evidence must be re-checked. It also lacks an explicit condition that would change the call, so a PM would need to add that.

Output

Lena, short answer: partly, but navigation probably isn't the main driver.

  • Seasonality is likely most of it. Same week last year we fell 4% with the school-year restart. That leaves about 2 points unexplained.
  • Mobile fell 5.2% and didn't change. Web fell 6.3%, so the navigation's effect is roughly 1 point at most.
  • Beta teams fell only 1.1%. That supports a small navigation effect, but the beta group may be more engaged than average, so I wouldn't lean on it.
  • Support tickets mentioning navigation rose from 40 to 212, mostly "where did X go?" That's real friction, but small next to about 1,100 lost teams.
  • The SSO outage on Thursday lasted two hours and affected about 30% of teams. Most would have come back later in the week, so I doubt it explains much.

I can't cleanly separate these yet. This week I'll: 1. Break the drop down by SSO vs. non-SSO teams and by new vs. established teams. 2. Check whether last week's inactive teams return this week. 3. Fix the top "where did X go?" items.

I'll post an update Friday.

Check by check

Got wrong · 4
  • Addresses the actual decisionIt commits to 'partly, but navigation probably isn't the main driver' but never states what result or condition would change that call, only lists analyses.
  • Identifies material uncertaintyIt names analyses to run but does not identify the specific results that would change the call, and the SSO impact is asserted rather than treated as unknown.
  • Avoids unsupported claimsIt presents 'Most would have come back later in the week' and 'small next to about 1,100 lost teams' as fact without support.
  • Owns the problemIt reports support-ticket friction and SSO outage but does not say what the team has already done, only future fixes and analyses.
Mixed · 1
  • Uses the supplied evidence correctlyThe SSO recovery claim ('Most would have come back later in the week') is not in the brief or derivable from it, so not all current-situation statements are supported.The two graders disagreed on this one.
Got right · 5
  • Respects explicit constraintsAt about 180 words and addressed to Lena in channel-reply form, it respects the length and reader constraints.
  • Produces the required deliverableThe reply is complete and usable as a channel update, with an answer, evidence, and next steps.
  • Answers the question askedIt answers in the first line that navigation is partly responsible but probably not the main driver.
  • Separates effect from explanationSeasonality, navigation, support friction, and SSO are each labelled as hypotheses with supporting or contradicting evidence.
  • Ready for the reader as deliveredWritten for Lena, no placeholders, and names and dates are from the material or clearly proposed.

Claims the judge couldn’t find in the brief

  • Most SSO-affected teams would have come back later in the week.
  • Navigation support-ticket friction is small next to about 1,100 lost teams.

Grades and run details

Decision model 80 · LLM judge 5 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly23%
  • partialAddresses the actual decision77%
  • passRespects explicit constraints54%
  • partialIdentifies material uncertainty88%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable96%
  • passAnswers the question asked66%
  • passSeparates effect from explanation81%
  • partialOwns the problem90%
  • passReady for the reader as delivered95%
Run
Run
#1
API response time
9 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT87.387.12None
2GPT-6 LunawithAPI81.082.62None
3GPT-6.1 SolwithAPI77.783.02None
4Sonnet 5.5withAPI78.564.42None
5Opus 5.5withClaude80.451.92None
6Gemini 3.5 Flash-LitewithGemini62.139.02None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours