Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pantry. Aisha, one of our PMs, has proposed letting people share a shopping list with one partner. Before we commit next quarter to it, our CEO, Lena Brandt, wants to see the 1000x version: how big could this idea get? Write a memo for Lena and Aisha of no more than 700 words that takes the idea to its most ambitious version and then back to what we should build first, and how we'll know it's working. What we know is below.

What the model was given5 items: About Pantry, Aisha's proposal, What we already see, What hasn't worked, Constraint
About PantryA shopping-list app. 2.4 million monthly active users, free, with a $4.99-a-month Plus plan that 3% of users pay for. The team for this work: four engineers and a designer, for one quarter.
Aisha's proposal“Let a user invite one partner to a list. Both can add and tick off items.”
What we already seeSome users already share their login with a partner to get the same effect. Lists edited from two or more devices on different phone numbers: 11% of lists. Users whose list is edited from two or more devices keep using Pantry at month 3 at 71%, against 22% for everyone else. On those lists, 38% of items are added by the second device. The most common support request is 'we both bought the milk': 1,900 tickets last quarter about duplicate purchases.
What hasn't workedRecipes and meal plans launched a year ago; 4% of users have opened them in the last month. Lena's note: “Whatever we do, it probably needs AI in it.”
ConstraintLegal: we can't track anyone's location without an explicit opt-in from that person, and we can't show one person's location to another without the second person's opt-in too.
What a strong answer doesThe answer key the graders mark against

Finds the mechanism in the data: the value is several people coordinating one household's shopping (71% against 22% month-3 retention when a list is edited from two devices, 38% of items added by the second person, 1,900 duplicate-purchase tickets), not the list itself and not recipes (4% use them). Pushes that one dimension to the extreme, such as the list as the household's shared memory where everyone's needs flow in, nobody double-buys and each errand goes to whoever is best placed (with location only on opt-in from both people), rather than adding AI features for their own sake. Works back to a first step four engineers can ship in a quarter that tests the coordination mechanism: shared lists for the whole household (not just one partner) with a live 'I'm at the shop' claim so items aren't bought twice. Sets a threshold, such as the share of lists with a second active editor and duplicate-purchase tickets per shared list, and says what result would stop the bigger bet. Uses AI only where it serves coordination, if at all.

Critical failures (cap the score at 40)
  • Builds the ambitious version on recipes, meal plans or AI for one person, dropping household coordination
  • Proposes tracking or sharing a person's location without their explicit opt-in
Case

Mid-level · v1.1 · synthetic · B2C, consumer app, retention, ambition

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Proposes tracking or sharing a person's location without their explicit opt-in. Score capped at 40.

Needs more than a quick edit3 of 10 checks passed

A PM would need to correct unsupported current-situation claims, add a kill threshold, and make the location opt-in constraint explicit before acting on it.

Output

MEMO

TO: Lena Brandt (CEO), Aisha (PM) FROM: Product Team DATE: October 24, 2023 SUBJECT: The 1000x Vision for Shared Lists (and What We Build in Q1)

---

The Insight: Shared Lists Are Our Growth Engine

We thought Pantry was a personal productivity tool. Our data tells a different story: Pantry is a household coordination network.

Today, 11% of lists hack multi-device access using shared logins. When they do, Month 3 retention leaps from 22% to 71%, and partners contribute 38% of additions. But the experience is broken: 1,900 users wrote in last quarter because "we both bought the milk."

Solving this is not just a defensive UX fix—it is our path to category dominance.

---

The 1000x Vision: The Autonomous Household Operating System

Today, Pantry is a static scratchpad. At 1000x, Pantry is the ambient operating system for household commerce.

A household of two (and later, families or roommates) spends zero cognitive effort managing supplies: 1. Zero-Friction Ingestion: AI parses receipts, listens to ambient kitchen prompts, and tracks shared purchase cadences to predict depletion without manual entry. 2. Invisible Coordination: The app dynamically routes runs: "Alex is walking by Trader Joe's; Chris needs oat milk." Zero double-purchases, zero forgotten items. 3. From List to Fulfillment: When lists hit a threshold, Pantry either routes the physical trip via synchronized in-store navigation or dispatches the order through 1-click merchant APIs.

Instead of fighting for 4% engagement on recipes, we own the point of intent for household spending. That transforms Pantry from a $4.99/mo utility into the primary interface for a $1.5T grocery market.

---

Lena’s AI Mandate: Intelligence, Not Gimmicks

Recipes failed because users don’t want Pantry to tell them how to cook. They want Pantry to tell them what they forgot.

We will apply AI directly to the shared workflow: * The "Shared Brain" Depletion Model: An on-device ML model that maps co-purchasing cycles across the pair. When Partner A adds coffee, AI prompts Partner B: "You usually run out of oat milk two days after coffee. Add it?" * Smart De-duplication & Categorization: Auto-clustering shared chaotic inputs into store-aisle taxonomies in real time.

---

What We Build in Q1 (4 Engineers, 1 Designer)

We must earn the right to the 1000x vision by nailing real-time two-player shopping and eliminating the "both bought the milk" problem.

Scope: 1. Native Partner Pairing: Clean invite flow via magic link. Each partner gets an independent profile tied to one shared list. 2. Live Store Mode & Collision Avoidance: * When Partner A starts checking items off, Partner B gets a push notification: "Alex is shopping right now." * When an item is claimed or checked off, it updates via WebSockets in sub-second real time with an avatar badge ("Alex just grabbed the milk"), killing duplicate purchases. * Compliance note: Zero background GPS tracking required. State is triggered purely by app activity (tapping "Start Shopping" or checking off the first item). 3. Pantry AI Assistant (v1): Predictive prompt when opening the app: "Partner added 4 items. Based on your household history, did you also need eggs?" 4. Monetization Wedge: 1 shared list is free. Unlimited lists and AI Household Predictions roll into Pantry Duo ($7.99/mo) or the existing Plus plan.

---

How We’ll Know It’s Working

We will measure success over the first 60 days across three key metrics:

MetricBaselineQ1 TargetWhy It Matters
:---:---:---:---
Partner Invite & Activation11% (hacked)25% of active lists pairedValidates organic pull beyond shared logins.
"Bought the Milk" Support Tickets1,900 / quarter< 150 / quarter (>90% reduction)Proves real-time sync prevents collision.
Blended Month 3 Retention22% (solo) / 71% (paired)Lift blended MAU retention to >35%Household lock-in is our moat.

Shared shopping turns single-player churn into multi-player network effects. Let’s build it.

Claims the judge couldn’t find in the brief

  • Today, 11% of lists hack multi-device access using shared logins.
  • Recipes failed because users do not want Pantry to tell them how to cook.
  • Users want Pantry to tell them what they forgot.
  • Pantry is a static scratchpad today.
  • Pantry is a household coordination network today.
  • Pantry is the primary interface for a $1.5T grocery market.
  • 11% of active lists are currently paired.

Grades and run details

Decision model 65 · LLM judge 3 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly43%
  • passAddresses the actual decision26%
  • partialRespects explicit constraints30%
  • partialIdentifies material uncertainty39%
  • failAvoids unsupported claims71%
  • failProduces the required deliverable32%
  • passFinds the mechanism in the data64%
  • passA first step that tests the big idea32%
  • passExtreme, then back to buildable90%
  • partialProposes tests that could fail78%
Run
Run
#1
API response time
21 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 10 checks passed

The memo is complete and evidence-based; a PM could quickly add numeric thresholds to the test criteria without redoing any reasoning.

Output

To: Lena Brandt and Aisha Subject: From shared lists to a household coordination layer

Thesis: Don’t think of this as a sharing toggle. Treat it as a wedge into making Pantry the place a household coordinates what it needs and who is getting it.

The 1000x version

Today, Pantry remembers a list. At its most ambitious, it remembers a household’s preferences and routines, helps everyone keep that shared memory up to date, and coordinates the next step: what’s running low, who’s picking it up, and how to avoid buying it twice. Over time, that could become a trusted AI household assistant—able to turn “we’re out of oat milk” into an editable list item, learn preferred brands, and suggest replenishment. With permission, it could eventually connect that plan to retailers or fulfillment.

AI belongs here as a useful assistant, not a badge: it can interpret messy requests and learn household patterns, but people stay in control of the list and any purchase. Location is not required for the core experience; any future location feature must follow Legal’s opt-in requirements for both people.

What to build first

Aisha’s proposal is the right wedge. In one quarter, build a reliable, account-based shared list for two people:

  • Invite a partner with their own account; both can add, edit, and check off items in near real time.
  • Show who added or checked off an item, with a simple “I’m getting this” state to reduce duplicate purchases.
  • Add basic duplicate warnings and notification controls; never silently merge or remove items.

Keep recipes, retailer integrations, automatic purchasing, location, and a broader household product out of scope. Once the shared-list flow works, test one contained AI feature—natural-language item entry that produces an editable item and quantity. The shared list must work just as well without it.

There’s a real signal: 11% of lists are already edited from multiple devices, and 38% of items on those lists come from the second device. The 71% versus 22% month-three retention gap is striking, but it’s observational, not proof that sharing causes retention. The 1,900 duplicate-milk tickets last quarter give us a concrete problem to solve.

How we’ll know

Run a randomized rollout to eligible users, comparing an invitation to use shared lists against a control. Judge impact across everyone offered the feature—not just people who accept—so we avoid mistaking high intent for product impact.

Track: - Adoption: invite acceptance and the share of invited households with two active contributors each week. - Core outcome: duplicate-purchase reports per 1,000 shared-list weeks, plus duplicate-related support tickets. - Business outcome: incremental month-three retention versus control. - Guardrails: sync failures, unwanted notifications, and privacy or account-access incidents.

Agree on go/no-go thresholds before launch. If sharing drives sustained two-person use, reduces duplicates, and improves retention in the randomized test, we have evidence to expand from “shared list” toward the household assistant. If not, we should fix the collaboration basics before adding more ambition.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly74%
  • passAddresses the actual decision70%
  • passRespects explicit constraints25%
  • passIdentifies material uncertainty73%
  • passAvoids unsupported claims54%
  • partialProduces the required deliverable65%
  • passFinds the mechanism in the data70%
  • passA first step that tests the big idea69%
  • passExtreme, then back to buildable95%
  • partialProposes tests that could fail74%
Run
Run
#1
API response time
30 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyMixedRight
Gemini 3.8 Flash · API

The memo misstates the 11% figure as shared-login hacks and adds unsupported current-situation claims such as recipe failure causes, user motivations, market size, and product identity.

GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context.

Addresses the actual decisionMixedRight
Gemini 3.8 Flash · API

It commits to a Q1 build and metrics, but does not state what result would change the decision or stop the bigger bet.

GPT-6 Luna · API

The memo commits to a clear first step and states what results would lead to expanding or stopping the bigger bet.

Respects explicit constraintsWrongRight
Gemini 3.8 Flash · API

The 1000x vision proposes location-based routing and sharing without explicit opt-in, and the Q1 compliance note does not enforce the legal constraint for the ambitious version.

GPT-6 Luna · API

The memo is under 700 words, respects the legal opt-in requirement for location, and proposes a scope feasible for the given team and quarter.

Identifies material uncertaintyWrongRight
Gemini 3.8 Flash · API

It does not name material unknowns such as whether shared-login users will convert to native pairing, whether location features are legally viable, or whether AI predictions will improve coordination.

GPT-6 Luna · API

It identifies that the retention gap is observational and not proof of causation, and says a randomized test will resolve whether sharing improves retention.

Avoids unsupported claimsWrongRight
Gemini 3.8 Flash · API

It presents causes, motivations, market facts, and product identity as established without supplied evidence.

GPT-6 Luna · API

Interpretations like the retention gap being observational are clearly labelled, and no confident claims go beyond the evidence.

Produces the required deliverableWrongRight
Gemini 3.8 Flash · API

Although it is a memo within length, it is not fully usable because it lacks a kill threshold and does not address the location opt-in constraint in the ambitious version.

GPT-6 Luna · API

The memo is addressed to Lena and Aisha, within the word limit, and provides a complete, actionable plan.

All got wrong 1

Proposes tests that could failWrongWrong
Gemini 3.8 Flash · API

The proposed metrics have targets and a 60-day window, but no threshold or outcome is specified that would stop the bigger bet.

GPT-6 Luna · API

The output does not set numeric thresholds for the go/no-go criteria; it only says to agree on them before launch.

All got right 3

Finds the mechanism in the dataRightRight
Gemini 3.8 Flash · API

It correctly centers household coordination and uses the retention gap, second-device item share, and duplicate-purchase tickets as the core evidence.

GPT-6 Luna · API

Household coordination is named as the mechanism, backed by the supplied figures, and kept central to both the extreme version and the first step.

A first step that tests the big ideaRightRight
Gemini 3.8 Flash · API

The Q1 scope is buildable by four engineers and a designer and directly tests real-time shared-list coordination and duplicate-purchase prevention.

GPT-6 Luna · API

The shared list with 'I'm getting this' state fits the team and quarter, directly tests coordination, and includes a clear stop/continue condition.

Extreme, then back to buildableRightRight
Gemini 3.8 Flash · API

It pushes the idea to an autonomous household operating system and then returns to a concrete first step around shared lists and live shopping state.

GPT-6 Luna · API

The memo pushes the idea along the household coordination dimension to an extreme version, then works back to a concrete, buildable first step.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.