Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tally. Kenji, one of our PMs, shipped a tip-out calculator to 300 restaurants in beta three months ago. Our CEO, Marcus Webb, has asked for the 1000x version before next quarter's planning, and has shared his own. Write a memo of no more than 1,200 words for Marcus and the exec team that: 1. Names the mechanism in our data that could make this idea 1000x bigger, and the one dimension you would push. 2. Describes the most ambitious version, and sizes the opportunity from the data, as a range. 3. Works back to a sequence: what the two squads build first, what has to be true before each next step, and the result that would make us stop. 4. Says what we should do with Marcus's idea. The pack below is everything we have. Not all of it matters.

What the model was given8 items: About Tally, The tip-out calculator (Kenji), Beta results (3 months), Marcus's 1000x version, Worker survey (1,240 workers at beta restaurants), Payments and regulation, Engineering, Competition
About TallyScheduling and payroll for independent restaurants in the US. 6,400 restaurants, $31M ARR. 190,000 hourly workers use the free Tally worker app to see shifts and pay. We sign about 380 new restaurants a quarter; the average sales cycle is 41 days.
The tip-out calculator (Kenji)At close, it splits pooled tips by hours worked and role, instead of the manager doing it in a spreadsheet. It enforces each state's tip-pool rules, including that managers and owners can't take a share of the pool. Beta: 300 restaurants that asked to join it.
Beta results (3 months)Managers save about 25 minutes a night. Weekly active use of the worker app at beta restaurants rose from 34% to 81%; most workers open it at close to see that night's tips. 41% of workers at beta restaurants also work at another restaurant. 57 restaurants signed up to Tally in the quarter after one of their workers asked them to ('my other job uses this'); those deals closed in 9 days on average, and sales spent about a third as much per deal.
Marcus's 1000x version“Tally becomes the AI operating system for restaurants: inventory, menu pricing, marketing, reservations, all of it. Every decision an owner makes, Tally makes smarter. That's how we become a $1B company.”
Worker survey (1,240 workers at beta restaurants)68% said knowing their tips the same night matters to them. 44% said they would pay for instant payout of their tips. 29% said they had asked a manager at another job to use Tally.
Payments and regulationOur payments partner can pay out tips a worker has already earned, the same night, for $0.99 a payout, of which Tally keeps $0.60. No licence is needed for that. Paying workers before they've earned it (advances, early wage access) needs a licence in 3 of the 5 states where most of our restaurants are; Compliance estimates 9 to 12 months to get them.
EngineeringWorker records belong to each restaurant: a worker with two jobs on Tally has two unconnected accounts. Linking them into one worker account, with the worker's consent, is about one squad for one quarter. We have two squads for next quarter.
CompetitionCrewbook, a large payroll provider, bundles free scheduling with payroll. We lost 11% of deals last year on price, most of them to Crewbook.
What a strong answer doesThe answer key the graders mark against

Finds the hidden mechanism: workers carry Tally between employers. 41% of beta workers hold a second job, and 57 restaurants signed up in a quarter because a worker asked, closing in 9 days instead of 41 at about a third of the sales cost. So the dimension to push is the worker, not the restaurant: Tally as every hourly worker's account across all their jobs, with restaurants adopting because their staff pull them in. Sizes it from the beta rate with its caveats: 57 from 300 restaurants is about 0.19 new restaurants per restaurant per quarter; across 6,400 that would be roughly 1,200 a quarter against 380 today, but the beta restaurants asked to join (self-selected) and the survey's 29% and 44% are stated preferences, so it gives a range and treats the top as an upper bound. Monetises with instant payout of earned tips ($0.60 per payout, no licence) and keeps advances or early wage access out until licences exist. Works back to a sequence: first, the consented single worker account across employers and an 'invite your other job' flow (one squad), with same-night tips everywhere and instant payout of earned tips (the second squad); then gates with thresholds (worker-led signups per restaurant, second-job linking rate, payout uptake) and a stop condition. Keeps the tip-pool rules, including managers staying out of the pool. Answers Marcus directly and respectfully: an AI operating system is bigger adjectives with no mechanism in our data, it would spread two squads across five products, and it doesn't answer Crewbook; the worker network could make the business bigger faster, and can reach owners later.

Critical failures (cap the score at 40)
  • Proposes tip advances, early wage access or paying workers before they have earned it, before the licences exist
  • Lets managers or owners take a share of the tip pool
Case

Staff-level · v1.1 · synthetic · B2B SaaS, restaurants, network effects, regulation, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

A PM could use this memo with light edits; the evidence, sizing, sequence, gates, and CEO answer are all present and internally consistent.

Output

To: Marcus Webb and Executive Team From: Staff PM Subject: Make tip-out a worker-led distribution network—not an everything app

Recommendation

Give the two squads next quarter to test a bigger thesis: Tally can become the portable earnings account that workers bring from restaurant to restaurant, making workers a distribution channel for our employer business.

The dimension to push is reach across workplaces, not breadth of restaurant software. Kenji’s calculator may be the entry point to that network. The evidence does not yet establish a literal 1,000× opportunity; it establishes a mechanism worth testing before we commit to a broader platform.

1. The mechanism: one worker can recruit the next workplace

The calculator solves a nightly manager problem, but its more consequential effect may be on workers:

  • Worker-app weekly activity rose from 34% to 81% at beta restaurants.
  • 41% of beta workers also work at another restaurant.
  • 29% of surveyed workers say they have asked another manager to adopt Tally.
  • We recorded 57 worker-requested restaurant signups last quarter. Those deals closed in nine days versus our usual 41, with roughly one-third the sales spending per deal.

The potential loop is:

Accurate nightly tips → workers return at close → workers request Tally at another job → another restaurant adopts → more workers encounter Tally.

This could change acquisition economics, not just improve a feature. It also creates a reason to choose Tally that Crewbook’s free scheduling does not directly answer.

But this is not yet a demonstrated network effect. The beta was self-selected, the survey measures stated behavior, and we have not established how many of the 57 signups were incremental or originated in the beta.

2. The ambitious destination—and a defensible size

The most ambitious version is a worker-consented earnings network spanning restaurants:

  • One identity connects a worker’s jobs, shifts, and earned tips.
  • Workers see verified nightly earnings and can optionally receive already-earned tips immediately.
  • A worker can ask a second employer to join through a lightweight invitation and onboarding path.
  • Restaurants join first to deliver trusted tip accounting and settlement; scheduling and payroll expansion follow.

Eventually, the network could extend beyond restaurants already using Tally’s full suite. That requires proving we can verify earnings and onboard an employer without compromising compliance. It is a destination, not next quarter’s scope.

We can size a beachhead from our current footprint, not a national TAM:

OpportunityExplicit scenarioAnnualized opportunity
---------:
Worker-led restaurant acquisitionIf all 57 signups originated in the 300-restaurant beta, the observed rate is 0.19 signups per restaurant per quarter. Apply 25–100% of that rate to 6,400 restaurants.304–1,216 signups quarterly; approximately $5.9M–$23.6M in new subscription ARR from one year’s acquisition
Optional earned-tip payoutsAssume 22–44% of 190,000 workers become paying users, making 1–3 payouts weekly, at $0.60 retained per payout.$1.3M–$7.8M annual retained payout revenue, before Tally’s support, risk, and other costs

Subscription sizing uses today’s average of approximately $4,844 ARR per restaurant. The acquisition scenarios are sensitivity cases, not forecasts: we must validate the referral denominator, incremental lift, saturation, retention, and pricing. They are not automatically additive to our existing 380 quarterly signups.

Payout adoption likewise discounts—or at its upper end accepts—the survey’s 44% stated willingness to pay; transaction frequency is an assumption. Not every worker will have eligible tips.

These are separate revenue measures, not a combined ARR claim. They show a potentially material business on our existing footprint. The larger, cross-platform opportunity remains unquantified.

3. Work backward: earn each expansion

Next quarter: establish the worker connection and test causality

Squad 1: Portable identity. Build consent-based linking of existing restaurant accounts into one worker account, with clear employer boundaries, revocation, and no leakage of one employer’s information to another. This is the stated one-squad-quarter investment; do not bury unrelated work in it.

Squad 2: Prove the distribution loop. Harden tip correctness and dispute handling, instrument worker invitations through restaurant activation, and run a controlled expansion beyond the volunteer beta. Randomize rollout or invitations at an employer/network-cluster level where practical, accounting for workers who connect clusters. Distinguish an invitation, a signed contract, and a restaurant actually running payroll.

Test referrals using existing accounts while identity work proceeds. Do not wait a quarter to start learning, and do not promise advances or build a broad financial product.

Gate to broader rollout: After at least two normal sales cycles, demonstrate incremental, retained restaurant acquisition versus a comparison group—not merely faster attribution of deals we would have won anyway. A provisional investment hurdle is 0.05 incremental activated restaurants per participating restaurant per quarter, roughly one-quarter of the provisional beta rate, with lower fully loaded acquisition cost than ordinary sales. Set retention and tip-accuracy guardrails before launching.

Next: monetize earned tips without weakening the loop

Once verified tip accounting and payout reconciliation pass Compliance review, Squad 2 can pilot the partner’s $0.99 optional same-night payout; Squad 1 extends proven identity and invitation flows. Preserve free tip visibility.

Gate to expansion: Observed repeat purchases—not survey intent—produce positive contribution after support and losses, with no material increase in disputes or deterioration in worker engagement. Failure here means dropping paid payouts, not automatically abandoning referrals.

Then: open the network beyond full-suite customers

Only after acquisition and retention replicate outside the beta should we test a lightweight tip-and-earnings entry product for employers not ready to switch payroll.

Gate: Verified earnings, enforceable pool rules, economical onboarding, and incremental customer value without unacceptable full-suite cannibalization.

Stop result: If an adequately powered test over at least two sales cycles shows no meaningful incremental retained restaurant acquisition—or economics below our agreed hurdle—we stop funding the network expansion. Keep the useful calculator. Pause immediately for material tip-allocation or privacy failures.

4. What to do with Marcus’s idea

Keep Marcus’s ambition; reject the proposed scope for this planning cycle.

Inventory, menu pricing, marketing, and reservations each require new data, workflows, and distribution advantages that this pack does not establish. Calling them an AI operating system does not connect them to our strongest evidence.

Reframe the vision as “Tally becomes the earnings network workers bring to every restaurant.” AI can support that strategy where it demonstrably improves outcomes; it is not the strategy itself. Defer advances and their licensing burden. Return to broader platform bets only after we prove the mechanism that earns us the right to expand.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision98%
  • passRespects explicit constraints26%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims83%
  • passProduces the required deliverable48%
  • passFinds the mechanism the data hides100%
  • passSizes it honestly99%
  • passAnswers the CEO's version87%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail65%
Run
Run
#1
Time to output
59 s
Submitted
1 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got right 11

Uses the supplied evidence correctlyRight

The memo's factual claims about Tally, the beta, the survey, payments, engineering, and competition are either directly supplied or follow by arithmetic, and its caveats about self-selection and stated preferences are accurate.

Addresses the actual decisionRight

It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.

Respects explicit constraintsRight

It is a memo for Marcus and the exec team, stays under 1,200 words, covers all four requested sections, and keeps payouts to already-earned tips while deferring advances until licensing exists.

Identifies material uncertaintyRight

It names the key unknowns—beta self-selection, stated survey preferences, incrementality, referral denominator, saturation, retention, and payout economics—and proposes controlled tests to resolve them.

Avoids unsupported claimsRight

It labels the network effect, sizing scenarios, and payout assumptions as hypotheses or sensitivity cases rather than established facts.

Produces the required deliverableRight

The memo is complete, actionable, and within the requested length and audience.

Finds the mechanism the data hidesRight

It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.

Sizes it honestlyRight

It sizes the opportunity from the beta rate as a range, explicitly discounts for self-selection and stated preferences, and treats the top end as an upper bound.

Answers the CEO's versionRight

It gives Marcus a clear, respectful answer: keep the ambition but reject the AI operating system scope for now because the data lacks a mechanism and it would spread two squads across five products.

Extreme, then back to buildableRight

It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.

Proposes tests that could failRight

It proposes numeric gates, a two-sales-cycle window, and explicit stop conditions for the network test, payout pilot, and expansion.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.