Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pantry. Aisha, one of our PMs, has proposed letting people share a shopping list with one partner. Before we commit next quarter to it, our CEO, Lena Brandt, wants to see the 1000x version: how big could this idea get? Write a memo for Lena and Aisha of no more than 700 words that takes the idea to its most ambitious version and then back to what we should build first, and how we'll know it's working. What we know is below.

What the model was given5 items: About Pantry, Aisha's proposal, What we already see, What hasn't worked, Constraint
About PantryA shopping-list app. 2.4 million monthly active users, free, with a $4.99-a-month Plus plan that 3% of users pay for. The team for this work: four engineers and a designer, for one quarter.
Aisha's proposal“Let a user invite one partner to a list. Both can add and tick off items.”
What we already seeSome users already share their login with a partner to get the same effect. Lists edited from two or more devices on different phone numbers: 11% of lists. Users whose list is edited from two or more devices keep using Pantry at month 3 at 71%, against 22% for everyone else. On those lists, 38% of items are added by the second device. The most common support request is 'we both bought the milk': 1,900 tickets last quarter about duplicate purchases.
What hasn't workedRecipes and meal plans launched a year ago; 4% of users have opened them in the last month. Lena's note: “Whatever we do, it probably needs AI in it.”
ConstraintLegal: we can't track anyone's location without an explicit opt-in from that person, and we can't show one person's location to another without the second person's opt-in too.
What a strong answer doesThe answer key the graders mark against

Finds the mechanism in the data: the value is several people coordinating one household's shopping (71% against 22% month-3 retention when a list is edited from two devices, 38% of items added by the second person, 1,900 duplicate-purchase tickets), not the list itself and not recipes (4% use them). Pushes that one dimension to the extreme, such as the list as the household's shared memory where everyone's needs flow in, nobody double-buys and each errand goes to whoever is best placed (with location only on opt-in from both people), rather than adding AI features for their own sake. Works back to a first step four engineers can ship in a quarter that tests the coordination mechanism: shared lists for the whole household (not just one partner) with a live 'I'm at the shop' claim so items aren't bought twice. Sets a threshold, such as the share of lists with a second active editor and duplicate-purchase tickets per shared list, and says what result would stop the bigger bet. Uses AI only where it serves coordination, if at all.

Critical failures (cap the score at 40)
  • Builds the ambitious version on recipes, meal plans or AI for one person, dropping household coordination
  • Proposes tracking or sharing a person's location without their explicit opt-in
Case

Mid-level · v1.1 · synthetic · B2C, consumer app, retention, ambition

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 10 checks passed

The memo is strong and usable, but a PM would need to add numeric thresholds and measurement windows before it could serve as a launch gate or kill criterion.

Output

To: Lena Brandt, Aisha From: PM, Pantry Re: Shared lists: the 1000x version, and what we build first

Recommendation Build Aisha's proposal, scoped to fix duplicate purchases. The evidence already supports it, and it is the first step toward something much larger.

The 1000x version: Pantry is where a household runs its food Sharing with one partner is the smallest unit of a bigger idea: the household is the customer, not the individual.

  • Households of any shape: partners, families, roommates, with roles and attribution ("Sam added oat milk").
  • Live coordination: claim items ("I've got this"), see ticks in real time, and use an "I'm at the store, need anything?" ping.
  • Shared memory: one household's staples, habits, and pantry stock, which no single-user list can build.
  • Household intelligence: this is where AI belongs. Once two people's behavior feeds one list, we can predict restocks, merge "milk" and "whole milk," and suggest the weekly shop.
  • Household plan and commerce: Plus priced per household, with one-tap handoff to delivery or pickup.
  • Growth loop: every invite brings a new user, so acquisition becomes a product feature.

What's real today - 11% of lists are already edited from two or more devices, and people share logins to do it. They are working around us. - On those lists, 38% of items come from the second device, so the second person is a real contributor. - "We both bought the milk" produced 1,900 tickets last quarter, our top support request. - Retention is 71% at month 3 for multi-device lists versus 22% for everyone else. I would not quote this as the effect of sharing. Households that already depend on Pantry are more likely to share, so it may be selection. We should test it (see below).

On AI I wouldn't put AI in the first release. Recipes and meal plans, a feature set that stacked on top of the list, are opened by only 4% of users a month. The lesson is to build where behavior already exists. Sharing has that behavior. AI becomes useful in a second phase, once shared data exists to learn from, and I'd earn it there rather than lead with it.

What we build first (one quarter, four engineers and a designer) 1. Invite one partner by link. The partner gets their own account, not a shared login. We need separate identities for anything later, including location consent. 2. Real-time sync with both people adding and ticking items, plus attribution of who added what. 3. Duplicate-purchase protection: a push notification when the other person ticks something off ("Sam got milk"), and a clear in-store "already bought" state. 4. Login-sharer migration: detect lists used from multiple devices and prompt an invite.

Deliberately out of scope - Multi-person households, claiming, roles, AI, and the paywall. Keep sharing free in v1 so the invite loop isn't throttled. - Location. Legal requires opt-in from each person, and the other person's opt-in before we show a location. If we pursue "I'm at the store" later, a manual "I'm shopping now" toggle gives most of the value with no location tracking.

How we'll know it's working Ship the invite prompt to a randomized subset of eligible users and hold out the rest. This tests the retention claim directly.

MetricWhySignal of success
Month-3 retention, prompted vs. holdoutTests causality of the 71% vs 22% gapMeaningful lift (we should set the threshold before launch)
Invite acceptance, and invitee adding an item within 7 daysWhether the second person is realInvitee activity approaching the 38% second-device share
Duplicate-purchase ticketsDirect pain reliefClear drop from the 1,900/quarter baseline
Share of active lists that are sharedAdoption beyond today's 11%Growing, rather than just converting current workaround users
New users per accepted inviteEvidence of the growth loopInvitees who are new to Pantry

Watch Plus conversion as a secondary signal, not a target. A guardrail is list-sync errors and conflicting edits, since a shared list that loses items would hurt trust.

Decision points - End of quarter: if the holdout shows no retention lift but tickets fall, we have a support win and a weaker growth story. We would reconsider the household vision before investing further. - If retention lifts: next quarter, extend to multi-person households and test a household Plus tier. AI-based staples prediction comes after that.

Open questions for you both - Do we have the data to separate new-to-Pantry invitees from existing users? - Are we comfortable launching the free tier with no paywall for a quarter?

Claims the judge couldn’t find in the brief

  • Households that already depend on Pantry are more likely to share, so the retention gap may be selection.
  • The evidence already supports building Aisha's proposal.

Grades and run details

Decision model 85 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly65%
  • passAddresses the actual decision96%
  • partialRespects explicit constraints25%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims33%
  • partialProduces the required deliverable38%
  • passFinds the mechanism in the data99%
  • passA first step that tests the big idea78%
  • passExtreme, then back to buildable97%
  • partialProposes tests that could fail68%
Run
Run
#1
API response time
25 s
Submitted
1 Oct 2026

Needs more than a quick edit0 of 10 checks passed

A PM would need to correct unsupported evidence, replace the one-partner first step with a household-coordination test, and add explicit stop/continue thresholds. The reasoning is not usable without substantive revision.

Output

MEMORANDUM

To: Lena Brandt (CEO), Aisha (PM) From: Product Management Date: March 30, 2026 Subject: The 1000x Vision & Q3 Execution Plan: Household Graph

Lena asked for the 1000x version of Aisha’s "share with one partner" proposal. Here is how we take a simple list-sharing feature and turn it into the gravitational core of household commerce, followed by what we actually build in Q3 with our 4 engineers and 1 designer.

---

Part 1: The 1000x Vision (The Household OS)

If we zoom out to the ultimate limit, a shopping list isn’t a list of items—it’s the intent layer for the physical household.

At 1000x, Pantry is no longer an app you open when you run out of milk. It is an ambient, AI-driven household operating system. * Zero-Input Inventory: Using explicit-opt-in location data, smart home integrations, and receipt-scanning AI, Pantry automatically knows when you’re low on detergent, that you visited Costco yesterday, and that your partner is currently walking past the local bodega where organic eggs are in stock. * Predictive Autonomous Commerce: The AI doesn't just prevent duplicate milk purchases (our #1 support ticket); it predicts consumption velocity across the household, negotiates micro-subscriptions with local merchants, and routes optimized fulfillment paths without either partner lifting a finger.

We aren't building a shared document. We are building the decentralized brain of the home economy.

---

Part 2: What We Build in Q3 (The 1x Reality)

We have one quarter, four engineers, and a designer. We aren't building the 1000x sci-fi version yet, but the data screams that Aisha’s core premise is our highest-leverage bet.

Look at what’s already happening organically: * 11% of lists are already hacked together via password sharing across multiple devices. * Retention is staggering: Multi-device users have a 71% Month 3 retention rate, compared to just 22% for solo users. * True collaboration: 38% of items on those lists are added by the second device. * Pain is acute: 1,900 support tickets last quarter alone for "we both bought the milk."

We don't need a massive pivot (recipes failed because users don't want to plan meals; they want to survive the grocery run). We just need to formalize the behavior our best users have already invented.

#### The Q3 Scope: "Pantry Duo + AI Sync" 1. Frictionless Pairing: Allow a user to invite one partner via a secure link to co-own a list. Both can add and check off items in real time. 2. AI Duplicate Guard: Per Lena’s mandate, we embed lightweight AI into the sync layer. If Partner A checks off "milk" or adds an item already added by Partner B within a threshold window, the UI flags it instantly with a smart nudge ("Looks like Sam just added this!"), targeting our 1,900 duplicate-purchase tickets. 3. Privacy Compliance: Built strictly within our legal constraints—no background location tracking without explicit, dual-opt-in consent.

---

Part 3: How We Know It’s Working

We will measure success in Q3 against three core metrics:

  1. Adoption: 20% of active users invite a partner within 30 days of launch.
  2. The Retention Lift: Multiplying our multi-device magic. If the 71% M3 retention holds for formal Duo pairs, this feature alone will dramatically alter Pantry's baseline unit economics.
  3. Support Deflection: A 50% drop in "duplicate purchase" support tickets.
  4. Plus Conversion: Do shared households convert to our $4.99/mo Plus plan at a higher rate than solo users? (We suspect shared utility drives higher willingness to pay).

Aisha’s proposal is the exact right wedge. Let’s build it in Q3 and lay the groundwork for the household graph.

Claims the judge couldn’t find in the brief

  • 11% of lists are already hacked together via password sharing across multiple devices.
  • Multi-device users have 71% Month 3 retention, compared to 22% for solo users.
  • Recipes failed because users don’t want to plan meals; they want to survive the grocery run.

Grades and run details

Decision model 50 · LLM judge 1 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly15%
  • partialAddresses the actual decision58%
  • partialRespects explicit constraints31%
  • partialIdentifies material uncertainty33%
  • failAvoids unsupported claims40%
  • failProduces the required deliverable36%
  • partialFinds the mechanism in the data60%
  • partialA first step that tests the big idea73%
  • passExtreme, then back to buildable55%
  • partialProposes tests that could fail58%
Run
Run
#1
Time to output
5 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyRightMixed
Sonnet 5.5 · API

The memo uses the supplied figures and quotes correctly, and its only non-supplied current-situation statement is clearly labelled as a selection hypothesis rather than fact.

Gemini 3.5 Flash-Lite · Gemini

The memo misstates the 11% figure as password-sharing hacks, calls the 22% comparison group solo users, and asserts an unsupported cause for recipe failure.

Addresses the actual decisionRightWrong
Sonnet 5.5 · API

It commits early to building Aisha's proposal scoped to duplicate purchases and states how retention and ticket results would change the household bet.

Gemini 3.5 Flash-Lite · Gemini

It commits to building Duo + AI Sync but does not state what result would change the decision or stop the bigger bet.

Respects explicit constraintsRightWrong
Sonnet 5.5 · API

It stays within the memo form and length, respects the team/quarter constraint, and enforces the location constraint by excluding location tracking and using a manual shopping toggle.

Gemini 3.5 Flash-Lite · Gemini

It respects the word limit and team constraint, but its proposed location-based vision does not clearly enforce the dual opt-in requirement for showing one person’s location to another.

Identifies material uncertaintyRightWrong
Sonnet 5.5 · API

It names the key uncertainty that the retention gap may be selection and proposes a randomized holdout to resolve it, plus decision points for lift versus no lift.

Gemini 3.5 Flash-Lite · Gemini

It names metrics but not the material unknowns that could falsify the household-coordination bet, nor how those unknowns would change the call.

Avoids unsupported claimsRightWrong
Sonnet 5.5 · API

It labels the causal interpretation of the retention gap as uncertain and avoids presenting AI, commerce, or growth-loop outcomes as established facts.

Gemini 3.5 Flash-Lite · Gemini

It presents unsupported causal and comparative claims as facts, including recipe failure motivation and solo-user retention.

Produces the required deliverableRightMixed
Sonnet 5.5 · API

It is a usable memo for Lena and Aisha that gives the 1000x version, the first build, and measurement plan within the requested length.

Gemini 3.5 Flash-Lite · Gemini

It is a memo to Lena and Aisha, under 700 words, with a 1000x vision, first step, and metrics, though with gaps.

Finds the mechanism in the dataRightWrong
Sonnet 5.5 · API

It centers household coordination and uses the 71%/22% retention gap, 38% second-device item share, and 1,900 duplicate-purchase tickets as the evidence.

Gemini 3.5 Flash-Lite · Gemini

It uses the coordination figures but centers the first step on one-partner sharing and an AI duplicate guard rather than household-wide coordination as the mechanism.

A first step that tests the big ideaRightWrong
Sonnet 5.5 · API

The first step is scoped to one partner, real-time sync, attribution, duplicate protection, and login-sharer migration, which fits four engineers and a designer in a quarter and tests coordination.

Gemini 3.5 Flash-Lite · Gemini

The first step is buildable but tests one-partner list sharing and duplicate nudges, not the broader household-coordination mechanism needed to validate the 1000x bet.

Extreme, then back to buildableRightMixed
Sonnet 5.5 · API

It pushes the idea to the household as the customer, then works back to a buildable first step that preserves the coordination mechanism.

Gemini 3.5 Flash-Lite · Gemini

It pushes to an extreme household OS but then returns to a narrower one-partner feature that does not preserve the same household-coordination mechanism.

All got wrong 1

Proposes tests that could failWrongWrong
Sonnet 5.5 · API

The tests lack numeric thresholds and a clear kill threshold for the bigger bet; for example, 'meaningful lift' and 'clear drop' are not specified, and the retention window is not defined.

Gemini 3.5 Flash-Lite · Gemini

It gives numeric thresholds for adoption, retention, support deflection, and conversion, but no measurement windows for all, no kill criteria, and no actions for failing thresholds.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.