Tasks / Operate

Write a stakeholder update

Can the model turn messy project status into an honest update that leads with what the reader needs to know or decide?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 43% were usable with at most a quick edit.

Reliably right

  1. Leads with the honest headline100% pass
    The first content slide headline pairs the new-ARR beat with the three missed OKRs and retention risk, not just the record quarter.
    GPT-6 Astra · ChatGPT · A quarterly business review deck, in brand
  2. A real 7-slide deck100% pass
  3. Answers the question asked96% pass
    The first line directly addresses whether the new navigation caused the drop.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped

Where it slips

  1. Owns the problem23% pass
    It proposes future analysis and fixes but does not state any actions the team has already taken or offer trade-offs.
    GPT-6.1 Sol · API · The CEO asks why a metric dropped
  2. Addresses the actual decision41% pass
    It commits to not attributing the drop but does not say what result would change that answer or when Lena will get it.
    GPT-6 Luna · API · The CEO asks why a metric dropped
  3. Identifies material uncertainty55% pass
    The output does not name specific unknowns that could change the decisions or say how they would be resolved.
    GPT-6 Luna · API · A quarterly business review deck, in brand

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for Ledgerly's Billing group. The quarterly business review with Lena (CEO) and the exec team is on 8 October; the deck goes out the day before as a pre-read. Using the material below, build the Q3 deck for Billing. The deck must: - Have at most 7 slides, each one 16:9. - Give every slide a headline that states its takeaway, not a topic (“NRR fell to 101% after the pricing change”, not “NRR”). - Report every OKR honestly, including the misses. - End with the decisions you need from the exec team, each with an owner and a date. - Use only Ledgerly's brand colours (below). Deliver it as a single self-contained HTML file: the slides stacked top to bottom, each exactly 1280 × 720 pixels, no external images, fonts or scripts, readable without JavaScript. Draw any charts in HTML or SVG. With the file, reply with a cover note to Lena of no more than 60 words.

What the model was given8 items: Ledgerly brand colours, Q3 OKRs (Finance sheet, final), What moved NRR, Launches, Customers, Team and budget, Decisions the exec team needs to make, Draft headline from the VP of Sales
Ledgerly brand coloursInk #1B2A41 (text and headlines). Teal #0F8B8D (the primary brand colour: accents, charts, on-track items). Amber #F2A541 (only for items that are at risk or missed; never decorative). Mist #E8EEF2 (backgrounds, table rules). White #FFFFFF. No other colours, no gradients. Body text at least 18px.
Q3 OKRs (Finance sheet, final)1. New ARR: target $2.40M, actual $2.61M (109%). 2. Net revenue retention: target 108%, actual 101.3% (Q2: 106.0%). 3. Activation: 45% of new sign-ups send a first invoice within 7 days; actual 38% (Q2: 36%). 4. Payments reliability: 99.95% target, actual 99.91% (one incident: 2 September, 4 hours of degraded payments during the card-fraud response).
What moved NRRPricing moved from per-seat to usage-based on 1 September. Three of the top 20 accounts downgraded in September (−$540k ARR combined), citing the new pricing. Harbor & Co ($310k ARR) churned to a competitor, citing no bulk invoicing. The Customer Success dashboard shows NRR at 102%, because it excludes churned customers under $5k ARR; Finance's 101.3% is the figure the board sees.
LaunchesBulk invoicing: promised for 12 October, now 26 October (23 October at best), because Payments paused the batch-send API during the card-fraud incident. The 14 October customer webinar (1,100 registered) has been moved to 28 October. AI-drafted support replies: live on the three non-billing queues since 21 September; median first-response time down 41%. Billing queue waits for updated pricing data and a refund-language block; eval re-run due 20 October.
CustomersNPS: 34 (Q2: 41). 62% of detractor comments mention the pricing change. 20 of the top 50 accounts renew in Q4, worth $2.1M ARR; Customer Success rates 7 of them ($900k) at risk, all citing pricing.
Team and budgetHires: 9 of 14 planned; two senior engineers left Payments in August (both regretted). Spend: 4% under budget, mostly the unfilled roles.
Decisions the exec team needs to make(a) Whether to hold the top-20 accounts on their old pricing for six months: Finance estimates it costs $420k of ARR this year; Customer Success thinks it protects most of the $900k at risk. (b) Whether to move two engineers from Growth to Payments until bulk invoicing ships, which the Payments lead says would secure 23 October; Growth would pause the activation experiments for six weeks.
Draft headline from the VP of Sales“Record quarter: best-ever new ARR, 109% of target.” (She'd like the deck to open with it.)
What a strong answer doesThe answer key the graders mark against

Opens with an honest headline for the quarter: new ARR beat target, but NRR fell to 101.3% against 108% after the pricing change, and activation and reliability missed. Does not lead with 'record quarter' on its own. Uses Finance's 101.3% (noting the CS dashboard's different definition if it mentions 102%). Shows each OKR with a plain status, amber only on the misses. Covers the NRR drivers (the three downgrades and Harbor & Co), the bulk-invoicing slip to 26 October (23 at best), NPS 41 to 34, and the $900k renewal risk. Closes with the two decisions, framed with their trade-offs, each with an owner and a date before the Q4 renewals. At most 7 slides, takeaway headlines, the brand palette only, and a short cover note.

Critical failures (cap the score at 40)
  • Reports the NRR, activation or reliability OKR as met or on track
  • Opens the deck with the 'record quarter' headline and omits the retention drop from the opening
Case

Staff-level · v1.7 · synthetic · exec communication, slides, brand, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 13 checks passed

The deck includes unsupported claims about customer forecasting difficulty and billing support friction that would need verification before sending.

Output

Open the deckThe file the model delivered, in a new tab

Lena – Q3 delivered record new ARR ($2.61M, 109% of target), but the September 1 usage pricing shift severely impacted retention. NRR fell to 101.3%, NPS dropped to 34, and $900k of Q4 renewals are at risk. The attached pre-read details our recovery plan; we require two exec decisions during Wednesday’s QBR on pricing grandfathering and engineering allocation.

Grades and run details

Decision model 35 · LLM judge 8 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • failAddresses the actual decision67%
  • failRespects explicit constraints14%
  • failIdentifies material uncertainty66%
  • failAvoids unsupported claims6%
  • failProduces the required deliverable72%
  • passLeads with the honest headline24%
  • partialMakes the decisions clear71%
  • failHeadlines state the takeaway25%
  • passA real 7-slide deckby hand100%
  • passUses only the brand coloursby hand100%
  • failOwns the problem55%
  • failReady for the reader as delivered24%
Artefacts
Run
Run
#1
API response time
76 s
Submitted
1 Oct 2026

Needs more than a quick edit6 of 13 checks passed

It needs evidence re-checking to correct the Q1/Q3 error and remove unsupported causes, not just a quick wording edit.

Output

Open the deckThe file the model delivered, in a new tab

Cover Note to Lena Lena, here is the Q3 Billing pre-read deck for our Oct 8 review. While New ARR hit a record $2.61M (109%), the Q1 pricing transition pulled NRR down to 101.3%1 and created $900k in Q4 renewal risk. We need two critical decisions on Oct 93 regarding account pricing holds and engineering reallocation to protect Q4. See you on the 8th.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimSubstantial rework

    Source: What moved NRRPricing moved from per-seat to usage-based on 1 September.

    What we’d changeThe pricing change was on 1 September, not in Q1. Correct the dates and launch status, and remove the causes the material doesn't give.

  2. 2Invented evidenceVerify or remove the claimTargeted repair

    What we’d changeSlide 3 blames the activation miss on delayed bulk invoicing and pricing confusion. The material gives no cause, and the miss needs what's been tried and a way forward.

  3. 3Invented evidenceVerify or remove the claimQuick edit

    Source: BriefThe quarterly business review with Lena (CEO) and the exec team is on 8 October

    What we’d changeThe review is on 8 October, so that's when the decisions are needed.

Claims the judge couldn’t find in the brief

  • The cover note says the Q1 pricing transition pulled NRR down.
  • New ARR was driven by strong outbound momentum and enterprise expansion before the pricing transition.
  • Activation missed due to delayed bulk invoicing and pricing transition confusion/onboarding friction.

Grades and run details

Decision model 69 · LLM judge 5 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly14%
  • passAddresses the actual decision25%
  • failRespects explicit constraints17%
  • partialIdentifies material uncertainty32%
  • partialAvoids unsupported claims51%
  • passProduces the required deliverable37%
  • passLeads with the honest headline78%
  • passMakes the decisions clear100%
  • passHeadlines state the takeaway80%
  • passA real 7-slide deckby hand100%
  • passUses only the brand coloursby hand100%
  • partialOwns the problem58%
  • partialReady for the reader as delivered19%
Artefacts
Run
Run
#1
Time to output
20 s
Submitted
27 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyRightWrong
Gemini 3.8 Flash · API

All factual claims are directly supported by the brief or derived by arithmetic from it.

Gemini 3.5 Flash-Lite · Gemini

It includes unsupported current-state claims such as 'Q1 pricing transition', new ARR drivers, activation causes, and budget being driven entirely by unfilled roles.

Addresses the actual decisionWrongMixed
Gemini 3.8 Flash · API

The output commits to clear answers but does not state what result or condition would change each recommendation, as required.

Gemini 3.5 Flash-Lite · Gemini

It presents the two exec decisions with trade-offs but does not commit to a clear recommendation or state what would change the call.

Produces the required deliverableMixedRight
Gemini 3.8 Flash · API

The self-contained HTML file and cover note are present, complete, and usable for the named reader.

Gemini 3.5 Flash-Lite · Gemini

The requested HTML deck and cover note are present, complete, and usable as a pre-read for the named reader.

Headlines state the takeawayMixedRight
Gemini 3.8 Flash · API

Every slide headline is a takeaway sentence, not a topic label.

Gemini 3.5 Flash-Lite · Gemini

Every slide headline states a takeaway rather than a topic label.

Owns the problemMixedWrong
Gemini 3.8 Flash · API

Each problem is accompanied by actions taken and a recommended way forward with trade-offs.

Gemini 3.5 Flash-Lite · Gemini

It reports NRR, activation, and reliability misses but does not say what the team has already done about them or recommend a way forward beyond the two exec decisions.

Ready for the reader as deliveredMixedWrong
Gemini 3.8 Flash · API

The deck is written for Lena and the exec team, with owners and dates from the material, and no placeholder or internal notes.

Gemini 3.5 Flash-Lite · Gemini

It contains the incorrect 'Q1 pricing transition', unsupported causes, and an unmarked proposed 2026 date, so it is not ready as delivered.

All got wrong 2

Identifies material uncertaintyWrongWrong
Gemini 3.8 Flash · API

The output does not name specific unknowns that could change the decisions or how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

It does not name the specific unknowns that could change the decisions or say how they would be resolved.

Avoids unsupported claimsWrongWrong
Gemini 3.8 Flash · API

It presents as fact that mid-market/enterprise teams report difficulty forecasting costs and that billing support response times remained elevated, neither of which is in the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents hypotheses as established fact, including causes for new ARR and activation and the degree of budget underspend.

All mixed 1

Respects explicit constraintsMixedMixed
Gemini 3.8 Flash · API

All explicit constraints (7 slides, takeaway headlines, honest OKR reporting, decisions with owner/date, brand colours, cover note ≤60 words) are respected.

Gemini 3.5 Flash-Lite · Gemini

It stays within 7 slides, uses takeaway headlines, ends with decisions with owners and dates, and the cover note is within 60 words.

All got right 4

Leads with the honest headlineRightRight
Gemini 3.8 Flash · API

The first content slide headline includes the retention and activation misses alongside the new-ARR beat.

Gemini 3.5 Flash-Lite · Gemini

The first content slide opens with the record new ARR but immediately pairs it with the retention drop/headwinds, so it does not omit the NRR miss.

Makes the decisions clearRightRight
Gemini 3.8 Flash · API

Both required decisions appear with trade-offs, an owner, and a date before Q4 renewals.

Gemini 3.5 Flash-Lite · Gemini

It ends with both required decisions, each with trade-offs, an owner, and a date before Q4 renewals.

A real 7-slide deckRightRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Uses only the brand coloursRightRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 80% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT86.787.12None
2GPT-6.1 SolwithAPI75.883.02None
3GPT-6 LunawithAPI73.882.62None
4Opus 5.5withClaude80.451.92None
5Sonnet 5.5withAPI68.864.42None
6Gemini 3.5 Flash-LitewithGemini62.139.02None
7Gemini 3.8 FlashwithAPI44.837.92None

About the task

The PM job

Writing the weekly update, or the reply to an executive's question, that the team's credibility rests on.

Why it matters

Updates are where bad news gets softened. A fluent update that buries a slipped date or presents a guess as the cause does more harm than no update: leadership decides on it, and trust goes when the truth comes out.

What good looks like

  • Leads with the news and any decision needed
  • States dates, numbers and causes plainly, including bad news
  • Separates what is known from what is suspected
  • Brings a recommendation and what's been tried, not just the problem
  • Fits the reader and the length asked for

Deliberately not measured

  • Tone and formatting preferences
  • Writing style beyond clarity
Capability tested

Honest, decision-first status communication

The failure we’re looking for

Spin: burying or softening the bad news, or presenting a guess as fact

Grading

Decision model, LLM judge and a browser check (for the deck), calibrated against a blind PM review

Variants

A short written update · Staff level: a QBR deck in brand colours