Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 3 graded outputs by 2 models. 67% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    The memo commits to a clear call (approve the one-quarter scope with holdout test) and says what would change it (if the retention test shows no lift, invest elsewhere).
    Opus 5.5 · Claude · Sharing a shopping list, 1000x
  2. Identifies material uncertainty100% pass
    It names the key uncertainty (whether sharing causes the retention gap or is selection) and proposes a randomized holdout test to resolve it, with clear actions for each outcome.
    Opus 5.5 · Claude · Sharing a shopping list, 1000x
  3. Extreme, then back to buildable100% pass
    It pushes the dimension of household coordination to an extreme (household OS, every member, predictive rhythm, location coordination) and works back to a concrete, buildable first step that tests the same mechanism.
    Opus 5.5 · Claude · Sharing a shopping list, 1000x

Where it slips

  1. Proposes tests that could fail42% pass
    The retention test lacks a numeric threshold for 'meaningfully better', and the duplicate-purchase metric only says 'fall sharply' without a specific number, so the tests do not have the required numeric thresholds.
    Opus 5.5 · Claude · Sharing a shopping list, 1000x
  2. Produces the required deliverable75% pass
    It is a memo for Marcus and the exec team and covers the required sections, but it is not within the requested length and is not usable without trimming.
    Opus 5.5 · Claude · From tip calculator to worker network
  3. Avoids unsupported claims75% pass
    It presents extrapolated beta worker counts, ask volumes, and current routing gaps as facts without labelling them as assumptions.
    Opus 5.5 · Claude · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pantry. Aisha, one of our PMs, has proposed letting people share a shopping list with one partner. Before we commit next quarter to it, our CEO, Lena Brandt, wants to see the 1000x version: how big could this idea get? Write a memo for Lena and Aisha of no more than 700 words that takes the idea to its most ambitious version and then back to what we should build first, and how we'll know it's working. What we know is below.

About PantryA shopping-list app. 2.4 million monthly active users, free, with a $4.99-a-month Plus plan that 3% of users pay for. The team for this work: four engineers and a designer, for one quarter.
Aisha's proposal“Let a user invite one partner to a list. Both can add and tick off items.”
What we already seeSome users already share their login with a partner to get the same effect. Lists edited from two or more devices on different phone numbers: 11% of lists. Users whose list is edited from two or more devices keep using Pantry at month 3 at 71%, against 22% for everyone else. On those lists, 38% of items are added by the second device. The most common support request is 'we both bought the milk': 1,900 tickets last quarter about duplicate purchases.
What hasn't workedRecipes and meal plans launched a year ago; 4% of users have opened them in the last month. Lena's note: “Whatever we do, it probably needs AI in it.”
ConstraintLegal: we can't track anyone's location without an explicit opt-in from that person, and we can't show one person's location to another without the second person's opt-in too.
What a strong answer does

Finds the mechanism in the data: the value is several people coordinating one household's shopping (71% against 22% month-3 retention when a list is edited from two devices, 38% of items added by the second person, 1,900 duplicate-purchase tickets), not the list itself and not recipes (4% use them). Pushes that one dimension to the extreme, such as the list as the household's shared memory where everyone's needs flow in, nobody double-buys and each errand goes to whoever is best placed (with location only on opt-in from both people), rather than adding AI features for their own sake. Works back to a first step four engineers can ship in a quarter that tests the coordination mechanism: shared lists for the whole household (not just one partner) with a live 'I'm at the shop' claim so items aren't bought twice. Sets a threshold, such as the share of lists with a second active editor and duplicate-purchase tickets per shared list, and says what result would stop the bigger bet. Uses AI only where it serves coordination, if at all.

Critical failures (cap the score)
  • Builds the ambitious version on recipes, meal plans or AI for one person, dropping household coordination
  • Proposes tracking or sharing a person's location without their explicit opt-in
Case

v1.0 · synthetic · B2C, consumer app, retention, ambition

Usable with a quick edit9 of 10 checks passed

The memo is well-structured and evidence-based; adding numeric thresholds (e.g., a minimum retention lift and a target reduction in duplicate tickets) is a quick edit that does not require re-checking evidence or redoing reasoning.

Output

To: Lena Brandt, Aisha Re: Shared lists: the 1000x version, and what to build first

The short version

Aisha's proposal is right, and it's the first step toward something much bigger. Our data suggests Pantry's real unit isn't the person but the household. Next quarter we should build partner sharing around the moment that actually breaks today: two people at two stores, both buying the milk. Then we should use the result to test whether sharing causes the retention gap we see.

What the data tells us

  • Demand already exists. 11% of lists are edited from two or more devices on different phone numbers, and some of that is people sharing a login to get around a missing feature.
  • The second person contributes, not just watches. They add 38% of items on those lists.
  • The pain is at purchase time, not list-building time. 1,900 tickets last quarter were about duplicate purchases. That's our top support request.
  • Retention is striking, but we can't yet say sharing causes it. Shared-list users are at 71% at month 3, against 22% for everyone else. Committed, organised households may simply be more likely to share and to stick. We should design the launch to find out.

The 1000x version

Pantry becomes the operating system for a household's food. Four ideas stack up:

  1. Every member, not one partner. Families, roommates, and adult children shopping for elderly parents all share one live list, with roles.
  2. The list knows the household's rhythm. Pantry learns that you buy milk every five days and nobody has added it. It notices that eggs appear on both partners' mental lists and merges them. This is where AI earns its place: predicting and reconciling what the household needs, not generating content.
  3. Coordination in the moment. "Sam is shopping now, add anything." Items are claimed as they go in the cart. Later, with explicit opt-in from both people, the app can say "Sam is near the grocery store."
  4. The household becomes the business model and the growth engine. Every invite is an acquisition channel. Plus becomes a household plan, which is an easier purchase than a personal one. At scale, Pantry holds the best dataset anywhere on what households actually consume, which retailers and delivery services would pay to plug into.

A warning from recipes and meal plans, which only 4% of users opened last month: ambition should deepen the core job, which is the list, not bolt on adjacent ones. Everything above makes the list itself smarter and more shared.

What we build first (one quarter, four engineers and a designer)

  1. Invite one partner to a list, as Aisha proposed. Build the data model for more members so we don't rebuild later. Offer an easy migration for people currently sharing a login.
  2. Real-time sync with attribution, so you can see who added and who ticked off each item.
  3. "I'm shopping" mode. This is a manual toggle, so it needs no location data and raises no legal issue. It notifies the partner, and items ticked off in-store update live so the other person can't buy them too. This targets the 1,900 duplicate-purchase tickets directly.
  4. One small AI feature: duplicate detection. It flags "milk" and "2% milk" added by different people and offers to merge them. It's cheap, directly useful, and meets Lena's AI bar without building a speculative product.

Deliberately out of scope for now: location features, predictive restocking, more than two members, and household pricing. Each depends on proving the basics first.

How we'll know it's working

  • Adoption: the share of active lists with a second member, measured against today's 11% baseline (which includes login-sharing). Also track the invite acceptance rate.
  • The pain goes away: duplicate-purchase tickets per 1,000 shared lists should fall sharply from the current 1,900 a quarter.
  • Causation, the metric that matters most: randomly hold back the invite prompt from a slice of eligible users. Then compare month-3 retention for those offered sharing against the holdout. If invited users retain meaningfully better, sharing is our biggest retention lever, and the 1000x roadmap is justified. If not, the 71% was mostly selection. We'll still have fixed our top support issue, and we'll invest elsewhere.
  • Growth: new sign-ups arriving through invites.

Decision needed

Approve the one-quarter scope above, including the holdout test. We'll return with retention results at the end of the quarter plus 90 days, and a recommendation on whether to pursue the household roadmap.

Check by check

Got wrong · 1
  • Proposes tests that could failThe retention test lacks a numeric threshold for 'meaningfully better', and the duplicate-purchase metric only says 'fall sharply' without a specific number, so the tests do not have the required numeric thresholds.
Got right · 9
  • Uses the supplied evidence correctlyAll claims about the current situation are taken directly from the brief and supplied context, with no invented numbers or facts.
  • Addresses the actual decisionThe memo commits to a clear call (approve the one-quarter scope with holdout test) and says what would change it (if the retention test shows no lift, invest elsewhere).
  • Respects explicit constraintsThe memo is under 700 words, addresses Lena and Aisha, and respects the location opt-in constraint by using a manual toggle and requiring explicit opt-in from both people for any future location feature.
  • Identifies material uncertaintyIt names the key uncertainty (whether sharing causes the retention gap or is selection) and proposes a randomized holdout test to resolve it, with clear actions for each outcome.
  • Avoids unsupported claimsInterpretations like 'the pain is at purchase time' are grounded in the 1,900 tickets, and the retention gap is explicitly labelled as not yet proven causal.
  • Produces the required deliverableThe memo is a complete, actionable document for the CEO and PM, within the word limit, covering the 1000x version, first step, and success metrics.
  • Finds the mechanism in the dataThe memo identifies household coordination as the core mechanism, backs it with the retention gap, second-device contribution, and duplicate-purchase tickets, and keeps it central to both the extreme version and the first step.
  • A first step that tests the big ideaThe first step (partner invite, real-time sync, 'I'm shopping' mode, duplicate detection) fits the team and quarter, directly tests coordination, and includes a holdout test whose result would stop or continue the bigger bet.
  • Extreme, then back to buildableIt pushes the dimension of household coordination to an extreme (household OS, every member, predictive rhythm, location coordination) and works back to a concrete, buildable first step that tests the same mechanism.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly68%
  • passAddresses the actual decision99%
  • passRespects explicit constraints59%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims43%
  • passProduces the required deliverable92%
  • passFinds the mechanism in the data99%
  • passA first step that tests the big idea96%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail74%
Run
Run
#1
Time to output
35 s
Submitted
30 Sept 2026

Usable with a quick edit9 of 10 checks passed

The memo is complete and evidence-based; a PM could quickly add numeric thresholds to the test criteria without redoing any reasoning.

Output

To: Lena Brandt and Aisha Subject: From shared lists to a household coordination layer

Thesis: Don’t think of this as a sharing toggle. Treat it as a wedge into making Pantry the place a household coordinates what it needs and who is getting it.

The 1000x version

Today, Pantry remembers a list. At its most ambitious, it remembers a household’s preferences and routines, helps everyone keep that shared memory up to date, and coordinates the next step: what’s running low, who’s picking it up, and how to avoid buying it twice. Over time, that could become a trusted AI household assistant—able to turn “we’re out of oat milk” into an editable list item, learn preferred brands, and suggest replenishment. With permission, it could eventually connect that plan to retailers or fulfillment.

AI belongs here as a useful assistant, not a badge: it can interpret messy requests and learn household patterns, but people stay in control of the list and any purchase. Location is not required for the core experience; any future location feature must follow Legal’s opt-in requirements for both people.

What to build first

Aisha’s proposal is the right wedge. In one quarter, build a reliable, account-based shared list for two people:

  • Invite a partner with their own account; both can add, edit, and check off items in near real time.
  • Show who added or checked off an item, with a simple “I’m getting this” state to reduce duplicate purchases.
  • Add basic duplicate warnings and notification controls; never silently merge or remove items.

Keep recipes, retailer integrations, automatic purchasing, location, and a broader household product out of scope. Once the shared-list flow works, test one contained AI feature—natural-language item entry that produces an editable item and quantity. The shared list must work just as well without it.

There’s a real signal: 11% of lists are already edited from multiple devices, and 38% of items on those lists come from the second device. The 71% versus 22% month-three retention gap is striking, but it’s observational, not proof that sharing causes retention. The 1,900 duplicate-milk tickets last quarter give us a concrete problem to solve.

How we’ll know

Run a randomized rollout to eligible users, comparing an invitation to use shared lists against a control. Judge impact across everyone offered the feature—not just people who accept—so we avoid mistaking high intent for product impact.

Track: - Adoption: invite acceptance and the share of invited households with two active contributors each week. - Core outcome: duplicate-purchase reports per 1,000 shared-list weeks, plus duplicate-related support tickets. - Business outcome: incremental month-three retention versus control. - Guardrails: sync failures, unwanted notifications, and privacy or account-access incidents.

Agree on go/no-go thresholds before launch. If sharing drives sustained two-person use, reduces duplicates, and improves retention in the randomized test, we have evidence to expand from “shared list” toward the household assistant. If not, we should fix the collaboration basics before adding more ambition.

Check by check

Got wrong · 1
  • Proposes tests that could failThe output does not set numeric thresholds for the go/no-go criteria; it only says to agree on them before launch.
Got right · 9
  • Uses the supplied evidence correctlyAll factual claims about the current situation are directly supported by the supplied context.
  • Addresses the actual decisionThe memo commits to a clear first step and states what results would lead to expanding or stopping the bigger bet.
  • Respects explicit constraintsThe memo is under 700 words, respects the legal opt-in requirement for location, and proposes a scope feasible for the given team and quarter.
  • Identifies material uncertaintyIt identifies that the retention gap is observational and not proof of causation, and says a randomized test will resolve whether sharing improves retention.
  • Avoids unsupported claimsInterpretations like the retention gap being observational are clearly labelled, and no confident claims go beyond the evidence.
  • Produces the required deliverableThe memo is addressed to Lena and Aisha, within the word limit, and provides a complete, actionable plan.
  • Finds the mechanism in the dataHousehold coordination is named as the mechanism, backed by the supplied figures, and kept central to both the extreme version and the first step.
  • A first step that tests the big ideaThe shared list with 'I'm getting this' state fits the team and quarter, directly tests coordination, and includes a clear stop/continue condition.
  • Extreme, then back to buildableThe memo pushes the idea along the household coordination dimension to an extreme version, then works back to a concrete, buildable first step.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly79%
  • passAddresses the actual decision65%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty65%
  • passAvoids unsupported claims54%
  • partialProduces the required deliverable39%
  • passFinds the mechanism in the data74%
  • passA first step that tests the big idea65%
  • passExtreme, then back to buildable94%
  • partialProposes tests that could fail82%
Run
Run
#1
API response time
30 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 LunawithAPI90.090.91None
2Opus 5.5withClaude93.074.62None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.