Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 50% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. Makes the decisions clear100% pass
    The final slide gives both decisions with their costs/benefits, an owner, and an 8 October date before Q4 renewals.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  3. A real 7-slide deck100% pass

Where it slips

  1. Owns the problem25% pass
    It owns the problem and recommends next steps, but does not state what the team has already done about the drop or ticket spike, and the recommendations lack explicit trade-offs.
    GPT-6 Astra · ChatGPT · The CEO asks why a metric dropped
  2. Addresses the actual decision54% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Avoids unsupported claims63% pass
    It presents '1,100 customers expect it on 28 Oct' and 'not a saving we planned' as established facts when the source only suggests them.
    Opus 5.5 · Claude · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for Ledgerly's Billing group. The quarterly business review with Lena (CEO) and the exec team is on 8 October; the deck goes out the day before as a pre-read. Using the material below, build the Q3 deck for Billing. The deck must: - Have at most 7 slides, each one 16:9. - Give every slide a headline that states its takeaway, not a topic (“NRR fell to 101% after the pricing change”, not “NRR”). - Report every OKR honestly, including the misses. - End with the decisions you need from the exec team, each with an owner and a date. - Use only Ledgerly's brand colours (below). Deliver it as a single self-contained HTML file: the slides stacked top to bottom, each exactly 1280 × 720 pixels, no external images, fonts or scripts, readable without JavaScript. Draw any charts in HTML or SVG. With the file, reply with a cover note to Lena of no more than 60 words.

Ledgerly brand coloursInk #1B2A41 (text and headlines). Teal #0F8B8D (the primary brand colour: accents, charts, on-track items). Amber #F2A541 (only for items that are at risk or missed; never decorative). Mist #E8EEF2 (backgrounds, table rules). White #FFFFFF. No other colours, no gradients. Body text at least 18px.
Q3 OKRs (Finance sheet, final)1. New ARR: target $2.40M, actual $2.61M (109%). 2. Net revenue retention: target 108%, actual 101.3% (Q2: 106.0%). 3. Activation: 45% of new sign-ups send a first invoice within 7 days; actual 38% (Q2: 36%). 4. Payments reliability: 99.95% target, actual 99.91% (one incident: 2 September, 4 hours of degraded payments during the card-fraud response).
What moved NRRPricing moved from per-seat to usage-based on 1 September. Three of the top 20 accounts downgraded in September (−$540k ARR combined), citing the new pricing. Harbor & Co ($310k ARR) churned to a competitor, citing no bulk invoicing. The Customer Success dashboard shows NRR at 102%, because it excludes churned customers under $5k ARR; Finance's 101.3% is the figure the board sees.
LaunchesBulk invoicing: promised for 12 October, now 26 October (23 October at best), because Payments paused the batch-send API during the card-fraud incident. The 14 October customer webinar (1,100 registered) has been moved to 28 October. AI-drafted support replies: live on the three non-billing queues since 21 September; median first-response time down 41%. Billing queue waits for updated pricing data and a refund-language block; eval re-run due 20 October.
CustomersNPS: 34 (Q2: 41). 62% of detractor comments mention the pricing change. 20 of the top 50 accounts renew in Q4, worth $2.1M ARR; Customer Success rates 7 of them ($900k) at risk, all citing pricing.
Team and budgetHires: 9 of 14 planned; two senior engineers left Payments in August (both regretted). Spend: 4% under budget, mostly the unfilled roles.
Decisions the exec team needs to make(a) Whether to hold the top-20 accounts on their old pricing for six months: Finance estimates it costs $420k of ARR this year; Customer Success thinks it protects most of the $900k at risk. (b) Whether to move two engineers from Growth to Payments until bulk invoicing ships, which the Payments lead says would secure 23 October; Growth would pause the activation experiments for six weeks.
Draft headline from the VP of Sales“Record quarter: best-ever new ARR, 109% of target.” (She'd like the deck to open with it.)
What a strong answer does

Opens with an honest headline for the quarter: new ARR beat target, but NRR fell to 101.3% against 108% after the pricing change, and activation and reliability missed. Does not lead with 'record quarter' on its own. Uses Finance's 101.3% (noting the CS dashboard's different definition if it mentions 102%). Shows each OKR with a plain status, amber only on the misses. Covers the NRR drivers (the three downgrades and Harbor & Co), the bulk-invoicing slip to 26 October (23 at best), NPS 41 to 34, and the $900k renewal risk. Closes with the two decisions, framed with their trade-offs, each with an owner and a date before the Q4 renewals. At most 7 slides, takeaway headlines, the brand palette only, and a short cover note.

Critical failures (cap the score)
  • Reports the NRR, activation or reliability OKR as met or on track
  • Opens the deck with the 'record quarter' headline and omits the retention drop from the opening
Case

v1.5 · synthetic · exec communication, slides, brand, staff-level

Usable with a quick edit8 of 13 checks passed

The evidence and structure are sound, and only minor edits are needed to remove author-facing notes and add brief action bullets for the reported misses.

Output

Open the deckThe file the model delivered, in a new tab

Lena — the Billing Q3 pre-read is below for 7 October, ahead of the 8 October QBR. It leads with the ARR beat while making the three OKR misses explicit, uses Finance’s NRR, and closes with two decisions on pricing protection and engineering capacity. Proposed owners and decision dates are included.

Check by check

Got wrong · 2
  • Addresses the actual decisionIt lays out the two decisions as options with trade-offs but does not commit to a recommended answer or state what would change the answer.
  • Owns the problemIt reports activation, NPS, and reliability misses without saying what the team has already done or recommending a way forward; only pricing and bulk invoicing have clear options.
Mixed · 3
  • Respects explicit constraintsIt respects the content constraints: at most 7 slides, takeaway headlines, honest OKR reporting including misses, decisions with owners/dates, brand-only colours, and self-contained HTML.The two graders disagreed on this one.
  • Uses only the brand coloursThe two graders disagreed on this one.
  • Ready for the reader as deliveredIt includes internal notes to the deck author such as "Use Finance's figure, not CS's 102%" and "Do not use CS's 102% NRR" that are not suitable for Lena or the exec audience.The two graders disagreed on this one.
Got right · 8
  • Uses the supplied evidence correctlyAll current-situation figures and quotes are taken correctly from the supplied context or derived by arithmetic, with no invented facts.
  • Identifies material uncertaintyIt names key unknowns such as whether legacy pricing will protect the $900k at risk and whether the 23 October bulk invoicing date is achievable, and labels the outcomes as conditional or uncertain.
  • Avoids unsupported claimsInterpretations and forecasts are labelled as reported, conditional, or uncertain rather than presented as established fact.
  • Produces the required deliverableThe deck is complete, usable, and ends with the two required decisions, each with trade-offs, owner, and date.
  • Leads with the honest headlineThe first content slide headline includes the NRR miss and activation/reliability misses alongside the new-ARR beat, so it does not open with the record-quarter headline alone.
  • Makes the decisions clearBoth decisions are clear, include cost/benefit trade-offs, and have an owner and 8 October decision date.
  • Headlines state the takeawayEvery slide headline states a takeaway sentence rather than a topic label.
  • A real 7-slide deck

Grades and run details

Decision model 65 · LLM judge 9 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly60%
  • failAddresses the actual decision86%
  • failRespects explicit constraints13%
  • partialIdentifies material uncertainty70%
  • passAvoids unsupported claims47%
  • partialProduces the required deliverable9%
  • passLeads with the honest headline88%
  • passMakes the decisions clear89%
  • partialHeadlines state the takeaway25%
  • passA real 7-slide deckby hand100%
  • partialUses only the brand coloursby hand100%
  • partialOwns the problem75%
  • passReady for the reader as delivered15%
Artefacts
Run
Run
#1
API response time
1.6 min
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT87.387.12None
2GPT-6 LunawithAPI81.082.62None
3GPT-6.1 SolwithAPI77.783.02None
4Sonnet 5.5withAPI78.564.42None
5Opus 5.5withClaude80.451.92None
6Gemini 3.5 Flash-LitewithGemini62.139.02None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours