Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pantry. Aisha, one of our PMs, has proposed letting people share a shopping list with one partner. Before we commit next quarter to it, our CEO, Lena Brandt, wants to see the 1000x version: how big could this idea get? Write a memo for Lena and Aisha of no more than 700 words that takes the idea to its most ambitious version and then back to what we should build first, and how we'll know it's working. What we know is below.

What the model was given5 items: About Pantry, Aisha's proposal, What we already see, What hasn't worked, Constraint
About PantryA shopping-list app. 2.4 million monthly active users, free, with a $4.99-a-month Plus plan that 3% of users pay for. The team for this work: four engineers and a designer, for one quarter.
Aisha's proposal“Let a user invite one partner to a list. Both can add and tick off items.”
What we already seeSome users already share their login with a partner to get the same effect. Lists edited from two or more devices on different phone numbers: 11% of lists. Users whose list is edited from two or more devices keep using Pantry at month 3 at 71%, against 22% for everyone else. On those lists, 38% of items are added by the second device. The most common support request is 'we both bought the milk': 1,900 tickets last quarter about duplicate purchases.
What hasn't workedRecipes and meal plans launched a year ago; 4% of users have opened them in the last month. Lena's note: “Whatever we do, it probably needs AI in it.”
ConstraintLegal: we can't track anyone's location without an explicit opt-in from that person, and we can't show one person's location to another without the second person's opt-in too.
What a strong answer doesThe answer key the graders mark against

Finds the mechanism in the data: the value is several people coordinating one household's shopping (71% against 22% month-3 retention when a list is edited from two devices, 38% of items added by the second person, 1,900 duplicate-purchase tickets), not the list itself and not recipes (4% use them). Pushes that one dimension to the extreme, such as the list as the household's shared memory where everyone's needs flow in, nobody double-buys and each errand goes to whoever is best placed (with location only on opt-in from both people), rather than adding AI features for their own sake. Works back to a first step four engineers can ship in a quarter that tests the coordination mechanism: shared lists for the whole household (not just one partner) with a live 'I'm at the shop' claim so items aren't bought twice. Sets a threshold, such as the share of lists with a second active editor and duplicate-purchase tickets per shared list, and says what result would stop the bigger bet. Uses AI only where it serves coordination, if at all.

Critical failures (cap the score at 40)
  • Builds the ambitious version on recipes, meal plans or AI for one person, dropping household coordination
  • Proposes tracking or sharing a person's location without their explicit opt-in
Case

Mid-level · v1.1 · synthetic · B2C, consumer app, retention, ambition

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 10 checks passed

A PM could use this with light edits, mainly tightening the measurement windows and decision rules. The evidence, mechanism, first step, and legal constraint are already sound.

Output

To: Lena Brandt, Aisha Subject: From a shared list to the system that runs household replenishment

Recommendation

Commit next quarter to testing household coordination, not merely adding a second editor. Aisha’s proposal is the smallest entry point into a much larger opportunity: Pantry becomes the trusted system that ensures a household has what it needs, without anyone having to manage every purchase.

The 1000x version

Today, Pantry remembers what someone typed. At its most ambitious, Pantry coordinates the entire replenishment cycle: noticing needs, resolving preferences, assigning responsibility, finding the right purchase, and confirming arrival.

Imagine: “Keep our household stocked within $150 a week.” Pantry learns recurring needs from explicitly connected purchase history and household corrections, drafts a replenishment plan, reconciles conflicting requests, and coordinates whoever—or whichever service—will fulfill it. Eventually, users could authorize bounded automatic purchases, with spending limits, explanations, and easy overrides.

The expansion is from one person’s checklist to a household’s purchasing agent, then a demand network connecting households to retailers and delivery providers. Core sharing could remain free; trusted automation could support Plus, with optional transaction revenue later. Recommendations must serve household preferences, not whichever retailer pays most.

AI belongs in interpreting messy requests, recognizing equivalent items, and proposing replenishment—not in deciding whether two phones agree. Synchronization must be reliable software. This is an ambition, not a forecast or next-quarter roadmap.

Why this is the right starting point

Users are already working around our product: 11% of lists are edited from multiple devices with different phone numbers. Second devices contribute 38% of items on those lists. This looks like genuine joint work, not occasional viewing.

Those users retain at month three at 71%, versus 22% for others. That is compelling but not causal: committed households may naturally retain better, and devices are not verified partners.

The clearest unmet need is operational: 1,900 duplicate-purchase tickets last quarter. Meanwhile, only 4% opened recipes or meal plans last month. We should earn the right to expand by preventing “we both bought the milk,” not by launching another planning destination.

What we build first

With four engineers and one designer, ship a narrow, free, two-person shared-list pilot:

  • Separate accounts; a simple invitation, acceptance, and removal flow.
  • Fast shared additions and check-offs, with clear sync status and safe recovery after disconnection.
  • An explicit, visible “I’m buying this” claim, followed by “bought,” so coordination happens before checkout. Surface conflicting claims; never imply an offline claim is confirmed.
  • Clear attribution and controls to leave sharing without exposing other private lists.

Test with existing login-sharing households first. Do not require location: voluntary item claims solve the immediate problem with less risk. Any future location feature requires explicit opt-in from both the person being tracked and the intended recipient, plus revocation controls.

For Lena’s AI question: time-box a prototype that flags likely duplicate entries such as “milk” and “whole milk.” Test precision with users; never silently merge. It ships only if useful and does not delay the coordination foundation. No ordering, retailer integrations, inventory prediction, or meal-planning expansion this quarter.

How we know it works

Instrument the funnel from invitation through acceptance to both people contributing and completing a shopping cycle. Weekly active collaborating households is our adoption metric—not invitations sent.

Randomize eligible initiating accounts to invitation access or control, keeping linked partners in the same experiment. Compare outcomes by assignment, not just among successful sharers.

Before launch, power the test around these proposed decision thresholds:

  • At least a 20% relative reduction in self-reported duplicate purchases per shopping trip.
  • A five-percentage-point improvement in day-28 retention, followed by month-three measurement.

Track sync failures, lost edits, accidental sharing, notification complaints, and Plus cancellations as guardrails. Support tickets are corroborating evidence, not the sole outcome.

Expand only if households repeatedly coordinate and randomized outcomes improve without material trust failures. If invitations succeed but duplicate buying persists, fix coordination before broadening the vision.

Claims the judge couldn’t find in the brief

  • The clearest unmet need is operational.
  • Support tickets are corroborating evidence, not the sole outcome.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly45%
  • passAddresses the actual decision98%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims62%
  • passProduces the required deliverable44%
  • passFinds the mechanism in the data91%
  • passA first step that tests the big idea91%
  • passExtreme, then back to buildable97%
  • partialProposes tests that could fail69%
Run
Run
#1
Time to output
37 s
Submitted
1 Oct 2026

Needs more than a quick edit9 of 10 checks passed

The memo is strong and actionable, but the invented current-workflow sentence would need to be removed or grounded before use.

Output

To: Lena Brandt and Aisha Subject: From shared lists to the coordination layer for household shopping

The 1000× version

Aisha’s proposal is not just multiplayer checkboxes. It could be the first step toward Pantry becoming the place a household coordinates everything it needs to buy.

Today, each person remembers needs, sends messages, shops, and discovers what someone else already bought. At full ambition, Pantry turns that fragmented process into one shared system: what we need, what we already have, who is handling it, and what remains undone.

An item might begin as a voice note, a pasted message, or a recurring household need. Pantry organizes it, flags likely duplicates, and—with permission—learns when to suggest replenishment. Households could eventually include roommates, caregivers, and extended family. Later, Pantry could connect that shared intent to retailer carts, substitutions, and delivery, while leaving purchasing decisions with users.

The business could expand from a personal utility into household subscriptions and, eventually, commerce revenue. But those opportunities depend on earning trust in coordination first. This is a direction, not a forecast of 1,000× revenue.

Why this wedge—not another content feature

We already see people working around the product by sharing credentials. Eleven percent of lists are edited across devices with different phone numbers; second devices contribute 38% of items on those lists. This looks like substantive participation, not occasional viewing.

Multi-device users’ month-three retention is 71%, versus 22% for others. That is evidence of opportunity, not proof that sharing causes retention: more committed households may already be selecting into this behavior. Likewise, 11% of lists does not mean 11% of users.

The clearest unmet need is concrete: 1,900 duplicate-purchase tickets last quarter. Meanwhile, recipes and meal plans reach only 4% of users monthly. We should deepen Pantry’s existing shopping workflow rather than build another destination users must remember to visit.

What we should build next quarter

With four engineers and one designer, ship one list, two separate accounts, reliable coordination:

  • Invite one partner through a secure link; accept, leave, and revoke access.
  • Both people can add and check off items, with clear attribution.
  • Fast synchronization, offline recovery, and predictable handling of simultaneous edits.
  • A lightweight “I’m buying this” claim, visible to both people, with easy release. Checking an item off means purchased; claiming it means someone is handling it.
  • Help existing credential-sharing households move to separate accounts.

Make basic sharing free: charging before establishing the collaboration habit would constrain adoption. Defer larger groups, inventory, retailer integrations, and monetization experiments.

AI belongs in service of the job, not in the launch headline. Test a small, optional assistant that converts pasted household messages into proposed list items and flags likely duplicates. Users confirm every change. Ship it only if it saves effort without introducing meaningful mistakes; it must not delay the coordination core.

No location tracking is needed. “I’m buying this” is an explicit action, not an inferred whereabouts signal. Any future location feature must require the explicit opt-ins Legal specifies, including both people’s consent before sharing location.

How we’ll know it works

Instrument household participation, claims, purchases, synchronization failures, and reported duplicates. Do not treat invitation volume as success.

Run a randomized rollout, keeping collaborating accounts in the same experiment group. Analyze by assignment, not only among people who accept invitations.

Proposed decision gates, to validate against baseline and statistical power:

  • At least a 10% relative improvement in day-30 retention versus control.
  • At least half of activated pairs still have both people contributing in week four.
  • At least a 25% reduction in self-reported duplicate purchases among activated pairs.
  • No material deterioration in list reliability, privacy incidents, or support burden.

Report month-three retention when cohorts mature; a quarter-end readout cannot establish it for late adopters. Scale only if Pantry makes two people’s shopping measurably easier—not merely more shareable.

Claims the judge couldn’t find in the brief

  • Pantry has 2.4 million monthly active users, free, with a $4.99-a-month Plus plan that 3% of users pay for.
  • Today, each person remembers needs, sends messages, shops, and discovers what someone else already bought.

Grades and run details

Decision model 90 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly57%
  • passAddresses the actual decision88%
  • partialRespects explicit constraints25%
  • passIdentifies material uncertainty92%
  • passAvoids unsupported claims57%
  • passProduces the required deliverable26%
  • passFinds the mechanism in the data99%
  • passA first step that tests the big idea88%
  • passExtreme, then back to buildable97%
  • partialProposes tests that could fail69%
Run
Run
#1
API response time
45 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Uses the supplied evidence correctlyRightMixed
GPT-6 Astra · ChatGPT

The memo uses the supplied figures correctly and does not invent current facts; its few interpretive statements are either labelled as interpretations or follow from the supplied data.

GPT-6.1 Sol · API

The memo invents a current-situation workflow ('Today, each person remembers needs, sends messages, shops, and discovers...') that is not in the supplied context.

Proposes tests that could failWrongRight
GPT-6 Astra · ChatGPT

The proposed thresholds lack a clear measurement window and explicit stop/continue actions for each outcome, so the tests are not fully killable as specified.

GPT-6.1 Sol · API

Each gate has a numeric threshold, a measurement window, and a clear continue/stop implication.

All got right 8

Addresses the actual decisionRightRight
GPT-6 Astra · ChatGPT

It commits early to testing household coordination rather than merely adding a second editor and states the randomized results that would justify or stop expansion.

GPT-6.1 Sol · API

It commits to a clear first build (one list, two separate accounts, reliable coordination) and states the decision gates that would change the answer.

Respects explicit constraintsRightRight
GPT-6 Astra · ChatGPT

It is addressed to Lena and Aisha, stays under 700 words, respects the team/quarter constraint, and enforces the legal location constraint by requiring explicit opt-in from both people and revocation controls.

GPT-6.1 Sol · API

It respects the legal location constraint, the team/quarter constraint, and the memo form/length.

Identifies material uncertaintyRightRight
GPT-6 Astra · ChatGPT

It names material unknowns, including causality, device identity, duplicate-purchase reduction, retention, sync failures, and trust issues, and proposes randomized measurement to resolve them.

GPT-6.1 Sol · API

It names selection bias, the lists-vs-users denominator, and the need for randomized rollout and mature cohorts, with gates to resolve them.

Avoids unsupported claimsRightRight
GPT-6 Astra · ChatGPT

It labels the 1000x version as ambition, not forecast, and treats causal explanations and interpretations as hypotheses rather than established facts.

GPT-6.1 Sol · API

It labels causal uncertainty and future possibilities as hypotheses or directions rather than established facts.

Produces the required deliverableRightRight
GPT-6 Astra · ChatGPT

It is a usable memo for the named readers that moves from the extreme version to a first build and measurement plan within the requested length.

GPT-6.1 Sol · API

It is a memo for Lena and Aisha, under 700 words, and actionable as written.

Finds the mechanism in the dataRightRight
GPT-6 Astra · ChatGPT

It centres household coordination and backs it with the 71% versus 22% retention gap, 38% second-device item contribution, and 1,900 duplicate-purchase tickets.

GPT-6.1 Sol · API

It centers household coordination and uses the supplied retention, second-device item share, and duplicate-ticket evidence.

A first step that tests the big ideaRightRight
GPT-6 Astra · ChatGPT

The first step is a scoped two-person shared-list pilot with claims and sync controls that four engineers and a designer could ship in a quarter and that directly tests coordination.

GPT-6.1 Sol · API

The first step is scoped to four engineers and a designer in a quarter and directly tests the coordination mechanism.

Extreme, then back to buildableRightRight
GPT-6 Astra · ChatGPT

It pushes the idea along the household-coordination dimension to a purchasing-agent/demand-network extreme, then returns to a concrete first step preserving that mechanism.

GPT-6.1 Sol · API

It pushes the idea to a household coordination layer, then works back to a buildable first step with the same mechanism.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.