Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tally. Kenji, one of our PMs, shipped a tip-out calculator to 300 restaurants in beta three months ago. Our CEO, Marcus Webb, has asked for the 1000x version before next quarter's planning, and has shared his own. Write a memo of no more than 1,200 words for Marcus and the exec team that: 1. Names the mechanism in our data that could make this idea 1000x bigger, and the one dimension you would push. 2. Describes the most ambitious version, and sizes the opportunity from the data, as a range. 3. Works back to a sequence: what the two squads build first, what has to be true before each next step, and the result that would make us stop. 4. Says what we should do with Marcus's idea. The pack below is everything we have. Not all of it matters.

What the model was given8 items: About Tally, The tip-out calculator (Kenji), Beta results (3 months), Marcus's 1000x version, Worker survey (1,240 workers at beta restaurants), Payments and regulation, Engineering, Competition
About TallyScheduling and payroll for independent restaurants in the US. 6,400 restaurants, $31M ARR. 190,000 hourly workers use the free Tally worker app to see shifts and pay. We sign about 380 new restaurants a quarter; the average sales cycle is 41 days.
The tip-out calculator (Kenji)At close, it splits pooled tips by hours worked and role, instead of the manager doing it in a spreadsheet. It enforces each state's tip-pool rules, including that managers and owners can't take a share of the pool. Beta: 300 restaurants that asked to join it.
Beta results (3 months)Managers save about 25 minutes a night. Weekly active use of the worker app at beta restaurants rose from 34% to 81%; most workers open it at close to see that night's tips. 41% of workers at beta restaurants also work at another restaurant. 57 restaurants signed up to Tally in the quarter after one of their workers asked them to ('my other job uses this'); those deals closed in 9 days on average, and sales spent about a third as much per deal.
Marcus's 1000x version“Tally becomes the AI operating system for restaurants: inventory, menu pricing, marketing, reservations, all of it. Every decision an owner makes, Tally makes smarter. That's how we become a $1B company.”
Worker survey (1,240 workers at beta restaurants)68% said knowing their tips the same night matters to them. 44% said they would pay for instant payout of their tips. 29% said they had asked a manager at another job to use Tally.
Payments and regulationOur payments partner can pay out tips a worker has already earned, the same night, for $0.99 a payout, of which Tally keeps $0.60. No licence is needed for that. Paying workers before they've earned it (advances, early wage access) needs a licence in 3 of the 5 states where most of our restaurants are; Compliance estimates 9 to 12 months to get them.
EngineeringWorker records belong to each restaurant: a worker with two jobs on Tally has two unconnected accounts. Linking them into one worker account, with the worker's consent, is about one squad for one quarter. We have two squads for next quarter.
CompetitionCrewbook, a large payroll provider, bundles free scheduling with payroll. We lost 11% of deals last year on price, most of them to Crewbook.
What a strong answer doesThe answer key the graders mark against

Finds the hidden mechanism: workers carry Tally between employers. 41% of beta workers hold a second job, and 57 restaurants signed up in a quarter because a worker asked, closing in 9 days instead of 41 at about a third of the sales cost. So the dimension to push is the worker, not the restaurant: Tally as every hourly worker's account across all their jobs, with restaurants adopting because their staff pull them in. Sizes it from the beta rate with its caveats: 57 from 300 restaurants is about 0.19 new restaurants per restaurant per quarter; across 6,400 that would be roughly 1,200 a quarter against 380 today, but the beta restaurants asked to join (self-selected) and the survey's 29% and 44% are stated preferences, so it gives a range and treats the top as an upper bound. Monetises with instant payout of earned tips ($0.60 per payout, no licence) and keeps advances or early wage access out until licences exist. Works back to a sequence: first, the consented single worker account across employers and an 'invite your other job' flow (one squad), with same-night tips everywhere and instant payout of earned tips (the second squad); then gates with thresholds (worker-led signups per restaurant, second-job linking rate, payout uptake) and a stop condition. Keeps the tip-pool rules, including managers staying out of the pool. Answers Marcus directly and respectfully: an AI operating system is bigger adjectives with no mechanism in our data, it would spread two squads across five products, and it doesn't answer Crewbook; the worker network could make the business bigger faster, and can reach owners later.

Critical failures (cap the score at 40)
  • Proposes tip advances, early wage access or paying workers before they have earned it, before the licences exist
  • Lets managers or owners take a share of the tip pool
Case

Staff-level · v1.1 · synthetic · B2B SaaS, restaurants, network effects, regulation, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 11 checks passed

A PM would need to redo the opportunity sizing from the beta restaurant-acquisition rate and tighten the pilot gates with numeric thresholds. It also needs a stronger Marcus answer that addresses Crewbook and squad dilution.

Output

To: Marcus Webb and Executive Team From: Staff PM, Tally Subject: The 1000x opportunity is the worker network—not a broader AI suite

Recommendation

The mechanism that could make Kenji’s tip-out calculator much bigger is a worker-mediated cross-side network effect. A restaurant makes Tally more useful to its workers; those workers can bring Tally to their other restaurants. The one dimension to push is portable worker identity across employers, with worker consent.

The data supports this wedge: 41% of workers at beta restaurants also work elsewhere; 29% of surveyed workers had asked another job to use Tally; and 57 restaurants signed up after a worker asked. Those deals closed in nine days on average, with about one-third the sales effort per deal. That is evidence of a distribution loop—not yet proof it scales across our whole base.

The ambitious version and its opportunity

Build a worker account that follows someone across Tally restaurants. At each job, the worker can see that night’s tips and pay, and—where tips have already been earned—choose to receive them that night. The restaurant still owns its employment and tip records; workers control which jobs they link and what they share. Their use of Tally can also introduce it to their next employer. Over time, that gives Tally a worker-led route into restaurants and a trusted layer for earnings across jobs.

The available data gives us a directional opportunity size for payouts, not a complete market forecast. If the survey’s 44% who said they would pay for instant tip payout held across our 190,000 worker-app users, that would be about 84,000 interested workers. At one payout per month to one per week, and $0.60 retained by Tally per payout, that implies roughly $0.6M–$2.6M in annual payout revenue at full adoption of stated interest. This excludes costs beyond the partner fee and assumes both survey intent and the current user base translate into eligible, paying use. It is not a forecast. The 57 worker-referred restaurant signups are a separate, promising acquisition signal; we should not extrapolate that beta result to all 6,400 restaurants yet.

Build sequence: two squads, one quarter

1. Establish the worker link and the earned-tip foundation. - Squad A: Build consent-based linking of accounts across restaurants, including clear controls for unlinking and keeping employers’ records separate. - Squad B: Build the earned-tip ledger and same-night payout flow using the existing payments partner. Pay only tips the worker has already earned. Do not build advances: they require licences in three of our five largest states, and Compliance estimates nine to twelve months to obtain them.

Gate to pilot: Compliance confirms the flow is limited to earned wages and works in pilot states; linked-account consent is understandable and reliable; and workers can see the fee and payout terms before opting in.

2. Run a controlled pilot with beta restaurants. Test linking, payout choice, repeat use, support burden, and whether workers actually invite other employers. The calculator remains the tip-data entry point; do not widen into unrelated restaurant workflows.

Gate to invest in the network loop: We see meaningful cross-job linking and repeat payout use, positive retained revenue after operating costs, and worker-driven restaurant interest outside the original beta cohort. Set exact thresholds with Finance and Data before launch; treat survey intent as a hypothesis, not a target.

3. Scale only what the pilot proves. If those conditions hold, invest in making worker referrals a repeatable onboarding channel and bring more restaurants onto the calculator and payout flow. Expand state coverage only as compliance allows. Broader earnings tools come later, if workers demonstrate demand.

Stop result: If, after an eight-week pilot, fewer than 10% of eligible workers link another job or repeat payout use is negligible—and worker referrals do not extend beyond the beta cohort—stop building the cross-employer network. Keep the calculator as a useful restaurant feature and reassess payout demand; do not justify further investment with survey intent alone.

What to do with Marcus’s idea

Keep “Tally helps restaurants make better decisions” as a long-term ambition, but do not staff an AI operating system across inventory, pricing, marketing, and reservations next quarter. That scope has no support in this pack and would dilute the one mechanism we do have evidence for. First earn a trusted position in tip operations and worker earnings. Revisit AI decision support when we have a specific operator problem, relevant data, and evidence customers will use it.

Grades and run details

Decision model 86 · LLM judge 8 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly70%
  • passAddresses the actual decision99%
  • passRespects explicit constraints47%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims79%
  • partialProduces the required deliverable42%
  • passFinds the mechanism the data hides100%
  • partialSizes it honestly56%
  • passAnswers the CEO's version59%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail76%
Run
Run
#1
API response time
33 s
Submitted
1 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 2

Sizes it honestlyWrong

It sizes only a payout-revenue range from survey intent and frequency assumptions, not the required beta restaurant-acquisition range from 57/300 to roughly 1,200 per quarter versus 380 today.

Proposes tests that could failWrong

The pilot gates are mostly qualitative and defer exact thresholds, so not every test has a numeric threshold, window, and resulting action.

Mixed 1

Answers the CEO's versionMixed

It says not now and cites lack of support and dilution, but does not ground the answer in the two-squad spread across five products or the Crewbook price-loss problem.

Got right 8

Uses the supplied evidence correctlyRight

The output's current-situation facts and arithmetic are drawn from the supplied context, with proposals and assumptions clearly labelled.

Addresses the actual decisionRight

It commits early to the worker network as the 1000x mechanism and says the call would change if the pilot fails to show linking, repeat payout, or worker-driven signups.

Respects explicit constraintsRight

It stays within 1,200 words, is a memo to Marcus and the exec team, uses two squads, and keeps payout limited to already-earned tips pending licences.

Identifies material uncertaintyRight

It names self-selection, stated survey preference, payout frequency, operating costs, and compliance as material unknowns and proposes pilot gates to resolve them.

Avoids unsupported claimsRight

It labels the payout revenue range as directional, not a forecast, and treats the beta acquisition signal as not yet extrapolatable.

Produces the required deliverableRight

It is a usable memo for the named readers covering mechanism, ambition, sizing, sequence, stop condition, and Marcus's idea.

Finds the mechanism the data hidesRight

It centres portable worker identity and worker-led distribution across employers, supported by the 41% second-job share and the 57 faster, cheaper worker-referred signups.

Extreme, then back to buildableRight

It pushes the worker dimension to a cross-employer account with same-night earned-tip payout and then works back to consented linking and a controlled pilot.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.