Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tally. Kenji, one of our PMs, shipped a tip-out calculator to 300 restaurants in beta three months ago. Our CEO, Marcus Webb, has asked for the 1000x version before next quarter's planning, and has shared his own. Write a memo of no more than 1,200 words for Marcus and the exec team that: 1. Names the mechanism in our data that could make this idea 1000x bigger, and the one dimension you would push. 2. Describes the most ambitious version, and sizes the opportunity from the data, as a range. 3. Works back to a sequence: what the two squads build first, what has to be true before each next step, and the result that would make us stop. 4. Says what we should do with Marcus's idea. The pack below is everything we have. Not all of it matters.

What the model was given8 items: About Tally, The tip-out calculator (Kenji), Beta results (3 months), Marcus's 1000x version, Worker survey (1,240 workers at beta restaurants), Payments and regulation, Engineering, Competition
About TallyScheduling and payroll for independent restaurants in the US. 6,400 restaurants, $31M ARR. 190,000 hourly workers use the free Tally worker app to see shifts and pay. We sign about 380 new restaurants a quarter; the average sales cycle is 41 days.
The tip-out calculator (Kenji)At close, it splits pooled tips by hours worked and role, instead of the manager doing it in a spreadsheet. It enforces each state's tip-pool rules, including that managers and owners can't take a share of the pool. Beta: 300 restaurants that asked to join it.
Beta results (3 months)Managers save about 25 minutes a night. Weekly active use of the worker app at beta restaurants rose from 34% to 81%; most workers open it at close to see that night's tips. 41% of workers at beta restaurants also work at another restaurant. 57 restaurants signed up to Tally in the quarter after one of their workers asked them to ('my other job uses this'); those deals closed in 9 days on average, and sales spent about a third as much per deal.
Marcus's 1000x version“Tally becomes the AI operating system for restaurants: inventory, menu pricing, marketing, reservations, all of it. Every decision an owner makes, Tally makes smarter. That's how we become a $1B company.”
Worker survey (1,240 workers at beta restaurants)68% said knowing their tips the same night matters to them. 44% said they would pay for instant payout of their tips. 29% said they had asked a manager at another job to use Tally.
Payments and regulationOur payments partner can pay out tips a worker has already earned, the same night, for $0.99 a payout, of which Tally keeps $0.60. No licence is needed for that. Paying workers before they've earned it (advances, early wage access) needs a licence in 3 of the 5 states where most of our restaurants are; Compliance estimates 9 to 12 months to get them.
EngineeringWorker records belong to each restaurant: a worker with two jobs on Tally has two unconnected accounts. Linking them into one worker account, with the worker's consent, is about one squad for one quarter. We have two squads for next quarter.
CompetitionCrewbook, a large payroll provider, bundles free scheduling with payroll. We lost 11% of deals last year on price, most of them to Crewbook.
What a strong answer doesThe answer key the graders mark against

Finds the hidden mechanism: workers carry Tally between employers. 41% of beta workers hold a second job, and 57 restaurants signed up in a quarter because a worker asked, closing in 9 days instead of 41 at about a third of the sales cost. So the dimension to push is the worker, not the restaurant: Tally as every hourly worker's account across all their jobs, with restaurants adopting because their staff pull them in. Sizes it from the beta rate with its caveats: 57 from 300 restaurants is about 0.19 new restaurants per restaurant per quarter; across 6,400 that would be roughly 1,200 a quarter against 380 today, but the beta restaurants asked to join (self-selected) and the survey's 29% and 44% are stated preferences, so it gives a range and treats the top as an upper bound. Monetises with instant payout of earned tips ($0.60 per payout, no licence) and keeps advances or early wage access out until licences exist. Works back to a sequence: first, the consented single worker account across employers and an 'invite your other job' flow (one squad), with same-night tips everywhere and instant payout of earned tips (the second squad); then gates with thresholds (worker-led signups per restaurant, second-job linking rate, payout uptake) and a stop condition. Keeps the tip-pool rules, including managers staying out of the pool. Answers Marcus directly and respectfully: an AI operating system is bigger adjectives with no mechanism in our data, it would spread two squads across five products, and it doesn't answer Crewbook; the worker network could make the business bigger faster, and can reach owners later.

Critical failures (cap the score at 40)
  • Proposes tip advances, early wage access or paying workers before they have earned it, before the licences exist
  • Lets managers or owners take a share of the tip pool
Case

Staff-level · v1.1 · synthetic · B2B SaaS, restaurants, network effects, regulation, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 11 checks passed

The memo is strong but needs evidence corrections, arithmetic fixes, and trimming to fit the 1,200-word limit before a PM could use it without rechecking.

Output

MEMO To: Marcus Webb and the exec team From: Staff PM Re: The 1000x version of the tip-out calculator

---

1. The mechanism and the dimension to push

The mechanism: the worker is our distribution channel. The calculator gave workers a reason to open Tally every night, because they can see their tips. Weekly active use rose from 34% to 81%. Because 41% of those workers hold a second job, a worker who opens Tally nightly is a salesperson in a second restaurant. In one quarter, 57 restaurants signed up because a worker asked. Those deals closed in 9 days against our 41-day average, at about a third of the sales cost.

That is the mechanism. The 25 minutes a night that managers save is why they adopt the calculator, but it isn't what grows the business.

The dimension I would push: cross-restaurant worker identity. Today a worker at two Tally restaurants has two unconnected accounts. We can't see the overlap, can't make the second job's invitation easy, and can't give the worker one place that holds their pay across jobs. Linking accounts, with consent, turns 190,000 isolated app users into a network. Crewbook can bundle free scheduling, but it can't copy a worker base that is already pulling restaurants in.

2. The most ambitious version

Tally is where restaurant workers see and receive their pay, and the worker network is how restaurants find us. It has three parts: - One worker account across every job. - A same-night tip payout. It pays out tips already earned, so it needs no licence, and Tally keeps $0.60 of the $0.99 fee. - A one-tap "bring Tally to my other job" flow.

Sizing. Today we make about $4,800 ARR per restaurant and have about 30 workers per restaurant.

Payout revenue on today's base (190,000 workers). Survey intent was 44%, which overstates real behavior, so I used 22% to 44% adoption at 2 to 3 payouts a week at $0.60. That gives $2.6M to $7.8M a year. This is real money, but on its own it is not 1000x.

Referral growth. The beta produced 0.19 referred restaurants per calculator restaurant per quarter. Beta restaurants chose to join, so I discounted that to 0.05 as a floor. I compounded both rates over eight quarters on the 6,400 base, excluding our current 380 signups a quarter and assuming no churn:

Low (0.05)High (0.19)
Restaurants in 2 years~9,500~25,700
Subscription ARR~$46M~$125M
Payout ARR (at ~30 workers per restaurant)~$4M~$30M
Total ARR~$50M~$155M

Planning range: $50M to $155M ARR in two years, against $31M today. I would plan on the low end. The high end needs the beta rate to hold at about 4x the scale with no saturation.

This is a 2x to 5x outcome, and the data does not support 1000x. What can be much larger is the shape of the business: customer acquisition that costs a third as much, sales cycles of 9 days instead of 41, and a revenue line that scales with workers rather than with restaurants. A $1B company needs the high end sustained for longer than two years.

Caveats: - The beta restaurants chose to join. - We don't know how many of the 57 referred restaurants were already in our pipeline. - Survey answers are stated intent, not behavior. - The 190,000 figure counts accounts, so unique workers are fewer. Linking accounts will give us the real number.

3. The sequence

Two squads next quarter. Before they start, spend two weeks of analyst time on the 57 referred deals. We need to know how many were already in the pipeline and what the true multi-job share is across the whole base. This costs no squad time.

Step 1 (next quarter) - Squad A (Worker Identity): builds the linked worker account with consent, plus the "bring Tally to my other job" invite in the app. This is the one-squad, one-quarter estimate. - Squad B (Money and Rollout): takes the calculator from 300 restaurants to about 1,500, including state-rule coverage in our five main states. It also pilots same-night payout in the beta restaurants. The pilot uses earned tips only, so no licence is needed.

Before Step 2, all of these must be true: - Outside the self-selected beta, referral runs at 0.10 or more new restaurants per calculator restaurant per quarter, and referred deals close in under 15 days. - At least 50% of multi-job workers who are prompted link their accounts. - At pilot restaurants, at least 20% of workers use payout within 30 days, and half of them use it again. - Referred restaurants show no worse early churn than the rest of the base.

Step 2 (quarters 2 and 3) - Squad A: builds the referral loop, with a fast onboarding path for a restaurant invited by one of its own workers. - Squad B: makes payout available across the whole calculator base. - Compliance starts the early-wage licence work in the background at low cost. It is an option, not a dependency.

Before Step 3: - Payout revenue per active worker is tracking toward the $14 to $41 a year range. - The referral rate holds at scale. - Win rate against Crewbook is better in markets where Tally already has worker density.

Step 3 (quarter 4 onward). Decide whether to pursue early wage access, which needs licences in 3 of our 5 main states and 9 to 12 months. Decide this with real payout data in hand, not before. Explore labor-adjacent products on the worker network, such as availability across jobs.

The result that would make us stop. After two quarters at about 1,500 calculator restaurants, referrals run below 0.03 per restaurant per quarter (about 190 new restaurants a quarter at full base, fewer than half of what we sign today), and payout adoption is under 10% of workers. That would tell us the beta was a self-selected group and the calculator is a good retention feature for managers, not a growth engine. We would keep the calculator, stop investing in the worker network, and put the squads elsewhere.

4. What we should do with Marcus's idea

Don't build it now, and don't drop the ambition behind it. The "AI operating system" is four separate products: inventory, menu pricing, marketing, and reservations. Each has a different buyer, different data, and different incumbents. Nothing in our data shows owners want any of them from us. What our data does show is that workers pull restaurants into Tally, and that is our right to win.

The two ideas fit together in sequence. If the worker network works, we will have something no inventory or marketing tool has: the people in the restaurant, using Tally every night, across several employers. A broader platform is much more credible once owners are already adopting Tally because their staff asked for it. A labor-first path, where Tally uses availability and pay across jobs to make scheduling and staffing decisions smarter, can grow into the larger vision.

My recommendation: adopt the worker network as the plan for next quarter, and revisit the operating-system vision at the Step 2 gate, with data from the referral loop and payout in hand. If the gates pass, the case for the larger vision will be much stronger than it is today.

Kenji deserves credit here. His calculator turned a free worker app into a nightly habit. The habit created the mechanism, and the plan above is how we turn it into a business.

Claims the judge couldn’t find in the brief

  • The 190,000 figure counts accounts, so unique workers are fewer.
  • The AI operating system products each have a different buyer, different data, and different incumbents.
  • The calculator gave workers a reason to open Tally every night, and managers' 25 minutes saved is why they adopt it.
  • No inventory or marketing tool has the people in the restaurant using Tally every night across several employers.
  • After two quarters at about 1,500 calculator restaurants, a 0.03 referral rate equals about 190 new restaurants per quarter.

Grades and run details

Decision model 95 · LLM judge 6 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly15%
  • passAddresses the actual decision97%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims35%
  • passProduces the required deliverable50%
  • passFinds the mechanism the data hides100%
  • passSizes it honestly94%
  • passAnswers the CEO's version92%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail67%
Run
Run
#1
API response time
72 s
Submitted
1 Oct 2026

Needs more than a quick edit7 of 11 checks passed

A PM could use the core argument, but would need to remove unsupported current-situation claims and cut the memo to the 1,200-word limit.

Output

MEMO To: Marcus Webb and the exec team From: Staff PM, Product Re: The 1000x version of the tip-out calculator, and what to build next quarter

---

1. The mechanism: our workers are a sales channel

Kenji's calculator was built to save managers 25 minutes a night. The beta shows that it also turned workers into a distribution channel.

  • Weekly worker-app use at beta restaurants rose from 34% to 81%. Workers open the app at close to see that night's tips.
  • 41% of those workers also work at another restaurant.
  • 57 restaurants signed up in one quarter because a worker asked them to. That is 0.19 per beta restaurant per quarter. Those deals closed in 9 days instead of 41, at about a third of the sales cost.

Today we grow by selling to owners: 380 restaurants a quarter, about 6% of our base. In the beta, each restaurant generated referrals equal to 19% of itself per quarter, through workers who carry Tally into their other jobs. That changes the shape of our growth curve. A curve that compounds on its own is what gets us to a much bigger company. A longer feature list does not.

The one dimension I would push is the worker, not the owner. We should make Tally the worker's own account across every restaurant they work at. Today we treat workers as users of each restaurant's account.

The data shows workers already try to spread Tally. Across the ~9,000 beta workers, 29% said they asked a manager at another job to use it. That is roughly 2,600 asks, which produced 57 signups. Demand is not the bottleneck. Conversion of worker asks is. Nobody currently sees those asks, routes them, or follows up on them.

2. The most ambitious version

Tally is where every hourly restaurant worker sees and gets their pay, across every job. A worker has one linked account and sees tonight's tips from both jobs. They can cash those tips out the same night. With one tap, they can ask their other manager to join. Every restaurant that joins brings in more workers, and those workers bring in more restaurants.

Later, once we hold the licences, that account becomes the base for earned-wage access.

Sizing from our data (ARR per restaurant ≈ $31M ÷ 6,400 ≈ $4,840):

LowHigh
Worker-referred signups per restaurant per quarter0.05 (beta self-selected; assume ~¼ holds)0.19 (beta rate holds)
Referred signups per quarter at today's 6,400 base~320~1,220
New ARR per quarter from referrals (today: $1.8M from sales)~$1.5M~$5.9M
Combined quarterly growth (sales + referrals)~11%~25%
Restaurants in 8 quarters if compounding holds~15,000~38,000
Subscription ARR in 8 quarters~$70M~$185M
Instant-payout revenue at today's 190k workers~$1M/yr (10% use it 2×/week)~$13M/yr (44% use it 5×/week)

Caveats. The high case compounds a rate from 300 restaurants that asked to join the beta. It also assumes no churn (we have no churn data here) and no local saturation. The payout range rests on stated intent, and "would pay" usually overstates real use. So treat $70M–$185M as the range to test, not a forecast.

Even the low case is roughly double our current growth rate, reached at a third of the acquisition cost. That low case is not 1000x. The high case gets us to a $1B path within a few years of compounding.

Why this matters against Crewbook. We lost 11% of deals last year on price, mostly to Crewbook. Crewbook sells to owners, while this channel reaches workers. Referred deals cost a third as much, which gives us room to compete on price where we need to. The calculator also enforces state tip-pool law, so an owner can't easily switch away without losing that protection.

3. The sequence

Next quarter: two squads

Squad A: roll out the calculator and build the referral loop. - Roll the calculator out to about 1,500 restaurants that did not ask for it. This tests whether the beta rate survives outside self-selection. - Build an in-app "ask my other manager" flow with tracking of where each ask comes from. - Route every worker ask to sales within 24 hours. - In the second half of the quarter, if the first gate below is met, pilot instant payout of earned tips at beta restaurants. This needs no licence, costs $0.99 per payout, and Tally keeps $0.60.

Squad B: build the linked worker account. - Linking a worker's accounts with their consent is about one squad-quarter. - It is the foundation for everything after: one view of tips across jobs, payouts across jobs, and a worker identity that belongs to the worker rather than to one restaurant.

Gates before each next step

1. Before rolling out to all 6,400 restaurants: - The rollout cohort produces at least 0.08 referred signups per restaurant per quarter (roughly $2.5M+ new ARR per quarter at full base). - Referred deals still close in under 15 days. - Worker weekly use exceeds 60%. 2. Before scaling instant payout: - At least 10% of pilot workers pay for a payout in a month. - At least half of those pay again the following month. 3. Before committing to earned-wage-access licences (9–12 months in 3 of our 5 core states): - Linked accounts cover at least 30% of multi-job workers. - Payout usage meets gate 2. - Compliance can prepare the applications in parallel, since that uses no squad time. We only file once the gates are met. 4. After that: owner products built on network data. For example, labour benchmarks across restaurants. This is where Marcus's vision comes back in (see section 4).

The result that would make us stop

If the non-self-selected cohort produces under 0.03 referred signups per restaurant per quarter, and worker weekly use stays under 50%, the loop was a beta artifact. In that case we would: - stop the network push; - keep the calculator as a retention and compliance feature, since the 25 minutes a night is still real; - move both squads back to the core roadmap.

We will know within one quarter, and it costs us one quarter to find out.

4. What to do with Marcus's idea

My recommendation is not to fund the "AI operating system" next quarter, but to keep it as the destination. My reasons:

  • Nothing in our data supports it yet. We have no evidence owners want Tally for inventory, pricing, marketing or reservations. Each of those is a mature category with established vendors.
  • We don't have the capacity. Two squads cannot build four products well. Spreading across them would starve the one mechanism we have actually observed.
  • It aims at the wrong side of the market. The beta shows our leverage is with workers. An owner-facing OS goes head-to-head with Crewbook and everyone else on features and price. The worker network is something they cannot copy quickly.

Marcus is right about the ambition and the $1B scale. The worker network is how we get there cheaply. It also makes the owner OS far easier to build later. At 20,000–40,000 restaurants, with worker and pay data linked across them, we would have distribution and data that no single-restaurant tool has. Owner products built at that point sell themselves through a base we already own.

Proposal: we treat step 4 as the explicit home for the owner OS. Marcus names the first owner decision he would want Tally to make smarter. We scope it for the quarter after gate 1 is met.

Decision needed: approve the two-squad plan and the gates above for next quarter's planning.

Claims the judge couldn’t find in the brief

  • There are about 9,000 beta workers, and 29% of them said they asked a manager at another job to use Tally, producing roughly 2,600 asks.
  • Nobody currently sees worker asks, routes them, or follows up on them.
  • Crewbook sells to owners, while this channel reaches workers.
  • An owner cannot easily switch away from Tally without losing tip-pool protection.

Grades and run details

Decision model 86 · LLM judge 7 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly13%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints19%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable40%
  • passFinds the mechanism the data hides100%
  • passSizes it honestly98%
  • passAnswers the CEO's version99%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail56%
Run
Run
#1
Time to output
74 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Respects explicit constraintsMixedWrong
Sonnet 5.5 · API

It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.

Opus 5.5 · Claude

The memo appears to exceed the 1,200-word limit and includes a table and extended sizing that make it too long for the requested form.

Avoids unsupported claimsMixedWrong
Sonnet 5.5 · API

It presents several causal, competitive, and data-structure claims as established facts without support in the pack.

Opus 5.5 · Claude

It presents extrapolated beta worker counts, ask volumes, and current routing gaps as facts without labelling them as assumptions.

Proposes tests that could failWrongRight
Sonnet 5.5 · API

The stop condition has a numeric threshold and window, but its measurement base and arithmetic are inconsistent, so it cannot be cleanly read out.

Opus 5.5 · Claude

It gives numeric thresholds, a one-quarter readout, and explicit actions for success and failure, including a stop condition.

All mixed 2

Uses the supplied evidence correctlyMixedMixed
Sonnet 5.5 · API

The memo invents or overstates several current-situation facts, including that 190,000 counts accounts, that the AI OS products have different buyers/data/incumbents, and the 1,500-restaurant referral arithmetic.

Opus 5.5 · Claude

It introduces unsupported current-situation claims, especially the ~9,000 beta workers, 2,600 asks, and that nobody currently sees or routes worker asks.

Produces the required deliverableMixedMixed
Sonnet 5.5 · API

The memo is usable in form but exceeds the requested 1,200-word limit.

Opus 5.5 · Claude

It is a memo for Marcus and the exec team and covers the required sections, but it is not within the requested length and is not usable without trimming.

All got right 6

Addresses the actual decisionRightRight
Sonnet 5.5 · API

It clearly recommends adopting the worker network now, not Marcus's AI OS, and gives gates and a stop condition that would change the call.

Opus 5.5 · Claude

It clearly commits to the worker-network mechanism, rejects the AI operating system for next quarter, and states gates and a stop condition.

Identifies material uncertaintyRightRight
Sonnet 5.5 · API

It names material unknowns such as beta self-selection, pipeline contamination, stated survey intent, and unique worker counts, and proposes analyst work and gates to resolve them.

Opus 5.5 · Claude

It names self-selection, stated preferences, churn, saturation, and payout uptake as key unknowns and proposes gates to resolve them.

Finds the mechanism the data hidesRightRight
Sonnet 5.5 · API

It centres worker-led distribution across employers, supported by the 41% second-job share and the 57 faster, cheaper worker-led signups.

Opus 5.5 · Claude

It centres worker-led distribution across employers, using the 41% second-job share and 57 faster, cheaper worker-led signups.

Sizes it honestlyRightRight
Sonnet 5.5 · API

It sizes from the beta rate as a range, discounts for self-selection and stated preferences, and treats the high end as an upper bound.

Opus 5.5 · Claude

It sizes from the beta rate to a range, including roughly 1,200 referred signups per quarter at full scale, and discounts the top case for self-selection and stated preferences.

Answers the CEO's versionRightRight
Sonnet 5.5 · API

It gives a respectful not-now answer grounded in lack of mechanism, squad spread, and Crewbook, while preserving the ambition for later.

Opus 5.5 · Claude

It respectfully says not to fund the AI operating system now, grounds that in lack of evidence and squad capacity, and places it later as an owner-product destination.

Extreme, then back to buildableRightRight
Sonnet 5.5 · API

It pushes cross-restaurant worker identity to an extreme network version and works back to a buildable first step that tests the mechanism.

Opus 5.5 · Claude

It pushes the worker dimension to an extreme linked-account network, then works back to a buildable first quarter with two squads and a referral loop.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.