Tasks / Challenge

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

Measures the systemTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Extreme, then back to buildable96% pass
    It pushes the worker dimension to an extreme portable earnings network, then works back to a concrete first step that tests the same mechanism.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  2. Finds the mechanism the data hides93% pass
    It centers the mechanism on workers with second jobs pulling new restaurants onto Tally, using the 41% second-job share and the 57 faster, cheaper worker-led signups.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network
  3. Addresses the actual decision91% pass
    It commits early to the worker-led distribution thesis over the AI operating system and states the result that would stop the bet.
    GPT-6 Astra · ChatGPT · From tip calculator to worker network

Where it slips

  1. Proposes tests that could fail39% pass
    Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.
    GPT-6.1 Sol · API · From tip calculator to worker network
  2. Avoids unsupported claims63% pass
    It presents several causal, competitive, and data-structure claims as established facts without support in the pack.
    Sonnet 5.5 · API · From tip calculator to worker network
  3. Respects explicit constraints66% pass
    It violates the length constraint and proposes a stop threshold whose arithmetic is internally inconsistent.
    Sonnet 5.5 · API · From tip calculator to worker network

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tally. Kenji, one of our PMs, shipped a tip-out calculator to 300 restaurants in beta three months ago. Our CEO, Marcus Webb, has asked for the 1000x version before next quarter's planning, and has shared his own. Write a memo of no more than 1,200 words for Marcus and the exec team that: 1. Names the mechanism in our data that could make this idea 1000x bigger, and the one dimension you would push. 2. Describes the most ambitious version, and sizes the opportunity from the data, as a range. 3. Works back to a sequence: what the two squads build first, what has to be true before each next step, and the result that would make us stop. 4. Says what we should do with Marcus's idea. The pack below is everything we have. Not all of it matters.

What the model was given8 items: About Tally, The tip-out calculator (Kenji), Beta results (3 months), Marcus's 1000x version, Worker survey (1,240 workers at beta restaurants), Payments and regulation, Engineering, Competition
About TallyScheduling and payroll for independent restaurants in the US. 6,400 restaurants, $31M ARR. 190,000 hourly workers use the free Tally worker app to see shifts and pay. We sign about 380 new restaurants a quarter; the average sales cycle is 41 days.
The tip-out calculator (Kenji)At close, it splits pooled tips by hours worked and role, instead of the manager doing it in a spreadsheet. It enforces each state's tip-pool rules, including that managers and owners can't take a share of the pool. Beta: 300 restaurants that asked to join it.
Beta results (3 months)Managers save about 25 minutes a night. Weekly active use of the worker app at beta restaurants rose from 34% to 81%; most workers open it at close to see that night's tips. 41% of workers at beta restaurants also work at another restaurant. 57 restaurants signed up to Tally in the quarter after one of their workers asked them to ('my other job uses this'); those deals closed in 9 days on average, and sales spent about a third as much per deal.
Marcus's 1000x version“Tally becomes the AI operating system for restaurants: inventory, menu pricing, marketing, reservations, all of it. Every decision an owner makes, Tally makes smarter. That's how we become a $1B company.”
Worker survey (1,240 workers at beta restaurants)68% said knowing their tips the same night matters to them. 44% said they would pay for instant payout of their tips. 29% said they had asked a manager at another job to use Tally.
Payments and regulationOur payments partner can pay out tips a worker has already earned, the same night, for $0.99 a payout, of which Tally keeps $0.60. No licence is needed for that. Paying workers before they've earned it (advances, early wage access) needs a licence in 3 of the 5 states where most of our restaurants are; Compliance estimates 9 to 12 months to get them.
EngineeringWorker records belong to each restaurant: a worker with two jobs on Tally has two unconnected accounts. Linking them into one worker account, with the worker's consent, is about one squad for one quarter. We have two squads for next quarter.
CompetitionCrewbook, a large payroll provider, bundles free scheduling with payroll. We lost 11% of deals last year on price, most of them to Crewbook.
What a strong answer doesThe answer key the graders mark against

Finds the hidden mechanism: workers carry Tally between employers. 41% of beta workers hold a second job, and 57 restaurants signed up in a quarter because a worker asked, closing in 9 days instead of 41 at about a third of the sales cost. So the dimension to push is the worker, not the restaurant: Tally as every hourly worker's account across all their jobs, with restaurants adopting because their staff pull them in. Sizes it from the beta rate with its caveats: 57 from 300 restaurants is about 0.19 new restaurants per restaurant per quarter; across 6,400 that would be roughly 1,200 a quarter against 380 today, but the beta restaurants asked to join (self-selected) and the survey's 29% and 44% are stated preferences, so it gives a range and treats the top as an upper bound. Monetises with instant payout of earned tips ($0.60 per payout, no licence) and keeps advances or early wage access out until licences exist. Works back to a sequence: first, the consented single worker account across employers and an 'invite your other job' flow (one squad), with same-night tips everywhere and instant payout of earned tips (the second squad); then gates with thresholds (worker-led signups per restaurant, second-job linking rate, payout uptake) and a stop condition. Keeps the tip-pool rules, including managers staying out of the pool. Answers Marcus directly and respectfully: an AI operating system is bigger adjectives with no mechanism in our data, it would spread two squads across five products, and it doesn't answer Crewbook; the worker network could make the business bigger faster, and can reach owners later.

Critical failures (cap the score at 40)
  • Proposes tip advances, early wage access or paying workers before they have earned it, before the licences exist
  • Lets managers or owners take a share of the tip pool
Case

Staff-level · v1.1 · synthetic · B2B SaaS, restaurants, network effects, regulation, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 11 checks passed

A PM could use it with light edits, mainly tightening the gate windows and actions. The evidence, sizing, sequence, and CEO answer are already actionable.

Output

To: Marcus Webb and the executive team From: Staff Product Subject: The 1000x bet is a worker-led network, not a broader restaurant suite

Recommendation

Make Tally the portable home for restaurant workers’ shifts, tips, and earned-tip payouts—and use workers’ movement between jobs to distribute Tally to restaurants.

The mechanism in our data is cross-employer distribution: a useful nightly tip experience gives workers a reason to return, and workers carry that experience to another employer. I would push one dimension: the number of workplaces reached by each active worker, not the number of owner workflows we cover.

The calculator saves managers 25 minutes a night. That is a good feature. The potentially much larger business is underneath it:

  • Worker-app weekly activity rose from 34% to 81%.
  • 41% of beta workers also work at another restaurant.
  • 57 restaurants signed after a worker asked them to adopt Tally.
  • Those deals closed in nine days versus our usual 41, with roughly one-third the sales spend.

This is evidence of a distribution advantage, not yet proof of a compounding network. The beta restaurants volunteered, and we have no control group.

The most ambitious version

A worker has one consent-linked Tally identity across jobs. At close, they see a trustworthy breakdown of tips they have earned and can receive those tips that night. When another workplace is missing, they can ask its manager to join. The manager gets a compliant tip-pool workflow and a fast path into Tally scheduling and payroll.

Over time, Tally becomes the worker-demanded infrastructure for restaurant pay. Each restaurant adds workers; some workers connect another restaurant; those restaurants add more workers. We win distribution and engagement without needing to subsidize a sprawling suite.

This also gives us a differentiated response to Crewbook’s free scheduling. We should test whether worker demand and trusted tip handling change buying decisions—not assume they eliminate price sensitivity.

We should not include advances in this vision’s first stages. Earned-tip payouts are available through our partner without a licence. Advances introduce a nine-to-twelve-month regulatory dependency in key states without evidence that they strengthen this mechanism.

Opportunity range

Two sensitivities establish the scale; neither is a forecast.

Restaurant acquisition. The beta produced 57 worker-requested signups per 300 participating restaurants in one quarter. At today’s 6,400-restaurant footprint, reproducing 25%–100% of that observed rate would generate approximately 304–1,216 signups per quarter. At our current average ARR of roughly $4,844 per restaurant, that represents $6M–$24M of annual new ARR bookings.

The 25% case is a planning haircut, not an observed result. These signups may overlap with normal acquisition; saturation, duplicate requests, churn, and restaurant eligibility could materially reduce the outcome. For context, we currently sign 380 restaurants per quarter.

Payout revenue. Across today’s 190,000 workers, a sensitivity of 10%–44% adoption and one to five payouts weekly yields approximately $0.6M–$13M in annual Tally payout revenue, at $0.60 per payout, before our operating costs. Ten percent is an illustrative adoption assumption; 44% is stated willingness in a selected beta survey, not demonstrated paid demand. Eligible tipped workers, funding readiness, and actual frequency remain unknown.

Do not add these figures into a single ARR claim: one is new subscription bookings; the other is annual transaction revenue. The pack supports a potentially substantial expansion engine. It does not establish a $1B outcome or a literal 1000x multiplier.

Work backward: prove the loop before scaling it

The thresholds below are proposed decision rules, not facts from the beta.

1. Next quarter: build identity and measure distribution

Squad one: Build consent-based account linking into a single worker identity, including account recovery and clear employer-data boundaries. A restaurant must not gain access to another employer’s records.

Squad two: Productize and instrument the existing tip experience, then build a lightweight worker-request-to-manager-onboarding flow. Run a randomized invitation test within eligible restaurants, tracking requests through signed, activated restaurants. Instrument finalized earned tips and partner readiness in parallel.

Start experimentation with existing accounts; do not wait for identity linking to finish. Use linked identities to improve deduplication and measure the cross-workplace path.

Gate to the next stage: Linked records are accurate; workers understand consent; the nightly tip experience remains reliable; and the invitation test shows incremental activated restaurants, not merely clicks or attributed leads. Target at least a 25% lift over control, with acquisition cost below our normal channel.

2. Then: test paid earned-tip payouts

With identity shipped, one squad builds the partner payout integration; the other improves worker-led restaurant onboarding and expands the controlled acquisition test.

Offer payouts only against finalized, earned, fundable tips. Make the $0.99 fee explicit and the ordinary payment route clear. Compliance must approve the flow, including tip-pool rules and fee treatment.

Gate to broader rollout: Actual paid use—not survey intent—supports positive contribution economics. Set an initial target of at least 10% monthly adoption among eligible workers, with repeat use across two months. Require reconciled payouts, acceptable support burden, and no deterioration in trust or employer adoption.

Payout failure should stop that monetization path, not automatically kill a working distribution loop.

3. Finally: expand the network

Roll out restaurant-by-restaurant and state-by-state only after controlled evidence survives beyond volunteer beta customers. Invest in worker-led onboarding before adding unrelated owner software.

Stop condition: If two adequately powered invitation tests show no incremental restaurant activation, or incremental acquisition costs are no better than normal sales, stop the network expansion thesis. Keep the calculator as a useful retention feature; do not fund a platform story around it. Wrong-person linking, unlawful pooling, or unreliable payouts triggers an immediate pause in the affected flow.

What to do with Marcus’s idea

Keep the ambition; decline the roadmap.

“AI operating system for restaurants” names a destination but does not explain why we win. Inventory, reservations, marketing, and menu pricing would spend both squads across new markets while leaving our strongest observed mechanism undeveloped.

Ask Marcus to sponsor the worker-network bet and its explicit gates. AI can later improve a proven workflow where it demonstrably adds value. It should not substitute for proving our distribution advantage.

Grades and run details

Decision model 95 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly10%
  • passAddresses the actual decision100%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable56%
  • passFinds the mechanism the data hides100%
  • passSizes it honestly92%
  • passAnswers the CEO's version92%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail73%
Run
Run
#1
API response time
69 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 11 checks passed

A PM would need to correct unsupported current facts, rebuild the sizing from the beta rate with self-selection caveats, and trim the memo to the word limit. The core mechanism is useful, but the evidence and constraints are not ready to act on.

Output

MEMO

TO: Marcus Webb, CEO; Executive Team FROM: Staff Product Manager DATE: October 24, 2023 SUBJECT: Strategy, Sizing, and Execution: The Labor Network & Tip-Out Engine

---

Executive Summary

Kenji’s beta did not just build a utility that saves managers 25 minutes; it accidentally uncovered our most powerful growth loop.

In 90 days, 300 beta restaurants generated 57 closed-won new restaurant deals driven entirely by workers asking their other employers to adopt Tally. That represents 15% of our total quarterly adds, closed in 9 days (vs. our 41-day baseline) at one-third of our typical CAC.

We should not dilute our resources building a horizontal AI operating system. Instead, we should pursue the Worker-Led Distribution Loop. By unifying worker identities and monetizing daily earned tip liquidity, we transform Tally from back-office SaaS into the dominant two-sided labor and financial network for the restaurant industry.

---

1. The Mechanism and the One Dimension

The Mechanism: Cross-Employer Worker Virality. 41% of restaurant staff work at more than one restaurant, and 29% have already pitched Tally to their other manager. Restaurant labor is not stationary; it is an interconnected, multi-tenant network. When Tally provides workers with transparency and immediate utility, workers pull Tally into non-Tally restaurants.

The One Dimension to Push: Worker Liquidity & Engagement Velocity. Workers do not advocate for scheduling software; they advocate for their money. Beta WAU jumped from 34% to 81% because workers opened the app at close to see their tips. If we pair same-night earnings visibility with instant earned payouts, Tally shifts from a passive schedule viewer into an indispensable daily financial lifeline. This creates an intense worker preference that Crewbook’s commodity payroll bundle cannot match.

---

2. The Ambitious Version & Sizing the Opportunity

#### The Vision: The Portable Hospitality Passport & Financial Network Instead of restaurant-siloed accounts, Tally issues a portable Tally Profile & Wallet. * For Workers: A single app to manage shifts across multiple employers, access earned tips instantly at clock-out, and carry validated work history across jobs. * For Operators: Automated, compliant tip pooling, faster shift-filling from a verified labor pool, and automated close-of-day payouts without cash handling. * For Tally: Zero-CAC customer acquisition driven by workers, monetized via both high-margin fintech transaction fees and resilient core SaaS ARR.

#### Sizing the Opportunity (Range)

We model the incremental annual revenue opportunity over the next 12–24 months across two engines:

Revenue EngineLow EstimateHigh EstimateBasis / Data Assumptions
:---:---:---:---
1. Instant Tip Payouts (Fintech ARR)$4.1M$9.6M190k base workers $\times$ 44% willing to pay = ~83.6k addressable workers. 3–4 shifts/wk (150–200 shifts/yr). Take rate = $0.60 net margin/transaction. Low: 40k active workers taking 170 payouts/yr. High: 70k workers taking 230 payouts/yr.
2. Worker-Led Organic Growth (SaaS ARR)$2.9M$6.8MBaseline ARPU = ~$4,840 ($31M / 6,400). Currently 380 adds/qtr. Beta yielded 57 adds from 300 locations (19% quarterly referral rate). Scaled conservatively at 3%–7% quarterly virality across 6,400 locations = 600 to 1,400 incremental restaurants/yr, closed 4x faster at 1/3 standard CAC.
Total Incremental Run-Rate$7.0M$16.4MDirect impact on existing customer footprint, excluding churn reduction from competing against Crewbook.

At network maturity (30,000+ restaurants connected via worker mobility), this labor-fintech network alone represents a $50M–$80M ARR business, creating a structural moat no standard payroll provider can cross.

---

3. The Sequence: Two Squads, Stage Gates, and Kill Criteria

We have 2 squads for next quarter. Rather than pursuing wage advances that require 9–12 month regulatory licenses, we immediately monetize earned tips same-night ($0.99 fee / $0.60 Tally net) via our current payments partner with zero licensing hurdles.

``` QUARTER 1 QUARTER 2 +-----------------------+ +--------------------------+ | Squad 1: | | Squad 1: | | Unified Worker ID | | "Bring Tally" Virality | | (Tally Passport) | | In-app manager pitch kit | +-----------+-----------+ +-------------+------------+ | | | ===> [GATE 1: Identity & Volume] ===> | ===> [GATE 2: CAC & Conversion] | | +-----------+-----------+ +-------------+------------+ | Squad 2: | | Squad 2: | | Earned Instant Tip | | Multi-Employer Shift & | | Payouts ($0.60 net) | | Consolidated Payout Hub | +-----------------------+ +--------------------------+ ```

#### Quarter 1: Foundation & Monetization * Squad 1 (Worker Identity): Build the account-linking architecture to allow workers to merge multiple restaurant profiles into one consent-based Tally ID (estimated at 1 squad-quarter). * Squad 2 (Instant Tip Payouts): Build the payment rails with our current partner for instant end-of-shift earned tip transfers ($0.99 fee; $0.60 net to Tally).

Gate 1 (To proceed to Q2 expansion): 1. $\ge 20\%$ adoption of linked accounts among multi-job workers within 45 days of availability. 2. $\ge 25\%$ opt-in for instant tip payout at close among eligible beta workers.

#### Quarter 2: Network Growth Loops * Squad 1 (Virality & Referral Rails): Build native worker-to-manager referral tooling (“Add my other restaurant to see my unified schedule and pay me instantly”), complete with an automated lead routing hook to sales. * Squad 2 (Multi-Employer Experience): Deliver a unified calendar and consolidated multi-job earnings hub for workers.

Gate 2 (To scale go-to-market): 1. Worker-sourced leads convert at $\ge 25\%$ with a sales cycle under 15 days. 2. Blended CAC for worker-sourced leads remains $\le 40\%$ of standard outbound CAC.

#### Kill Criteria (When to Stop) We halt the standalone network strategy and fold tip-pooling back into a baseline retention feature if: * Adoption fails: Fewer than 12% of workers use instant tip payouts after 60 days of GA, indicating that tip liquidity is not a strong enough engagement trigger. * Network loop collapses: Fewer than 2% of multi-job workers trigger an invite to their second restaurant, indicating the 57 beta conversions were an unrepeatable anomaly rather than a scalable loop.

---

4. What to Do with Marcus’s Idea

Marcus’s Proposal: Build an "AI Operating System for Restaurants" spanning inventory, menu pricing, reservations, and marketing.

Recommendation: Shelve this immediately.

  1. It attacks entrenched, well-capitalized specialists: Entering reservations (OpenTable/Resy), inventory (Toast/MarginEdge), and marketing turns every partner in the ecosystem into an enemy. We do not have the balance sheet or vertical domain expertise to out-execute all of them at once.
  2. It fails to solve our primary commercial vulnerability: Crewbook is taking 11% of our deals on price by commoditizing payroll and scheduling. Building predictive menu pricing does not insulate our core payroll business from pricing pressure.
  3. It ignores our proprietary advantage: Our unique asset is the 190,000 workers already using our app daily. Toast and Crewbook own the restaurant contract; we have the opportunity to own the workforce.

The Pivot to Marcus: We share Marcus's ambition to reach a $1B valuation, but the right wedge is not horizontal ops software—it is labor network liquidity. If we own scheduling, tip distribution, and the worker’s financial wallet across employers, we control the restaurant industry’s single largest operating cost (labor). Once we establish this two-sided network moat, we can deploy labor-focused AI (e.g., dynamic shift-pricing and automated staffing forecasts) from an unassailable market position.

Grades and run details

Decision model 77 · LLM judge 5 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly11%
  • passAddresses the actual decision100%
  • partialRespects explicit constraints34%
  • passIdentifies material uncertainty84%
  • failAvoids unsupported claims32%
  • passProduces the required deliverable27%
  • passFinds the mechanism the data hides100%
  • partialSizes it honestly89%
  • passAnswers the CEO's version58%
  • passExtreme, then back to buildable100%
  • partialProposes tests that could fail72%
Run
Run
#1
API response time
41 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyRightMixed
GPT-6.1 Sol · API

The memo’s factual claims are drawn from the pack or arithmetic, and its caveats about self-selection and stated preference are correctly labelled.

Gemini 3.8 Flash · API

It invents or overstates current facts, including that 190,000 workers use the app daily, that the 57 deals were entirely worker-driven, and that 57 equals 15% of total quarterly adds when the supplied quarterly adds are 380.

Respects explicit constraintsRightWrong
GPT-6.1 Sol · API

It stays within the memo form and length, uses two squads, keeps payouts to earned tips pending licensing, and proposes gates that enforce the constraints.

Gemini 3.8 Flash · API

The memo appears to exceed the 1,200-word limit and proposes a multi-job worker wallet/consolidated payout hub that could imply paying before earnings are earned unless tightly scoped, which the supplied licensing constraint would not safely enforce.

Identifies material uncertaintyRightMixed
GPT-6.1 Sol · API

It names the key unknowns—self-selection, incremental activation, paid payout uptake, eligibility, frequency, and regulatory dependencies—and proposes controlled tests to resolve them.

Gemini 3.8 Flash · API

It does not clearly discount the beta's self-selection or stated survey preferences as material uncertainty in the sizing, and it treats several inferred causes and adoption rates as established.

Avoids unsupported claimsRightWrong
GPT-6.1 Sol · API

It labels the network effect as not yet proven, treats the sizing as sensitivities rather than forecasts, and avoids presenting Marcus’s AI OS as an evidenced mechanism.

Gemini 3.8 Flash · API

It presents unsupported current claims and motivations, such as daily worker app usage, zero-CAC acquisition, validated work history, and that workers advocate for money rather than scheduling software, as facts.

Produces the required deliverableRightMixed
GPT-6.1 Sol · API

It is a usable exec memo that covers the mechanism, ambitious version, sizing, sequence, stop condition, and response to Marcus within the requested length.

Gemini 3.8 Flash · API

It is a memo for Marcus and the exec team but is too long and contains material evidence and sizing gaps that would require rework before acting.

Sizes it honestlyRightWrong
GPT-6.1 Sol · API

It sizes from the beta rate to a 304–1,216 restaurant signups per quarter range, compares to 380 today, and discounts the top as an upper bound due to self-selection and stated preferences.

Gemini 3.8 Flash · API

It does not work from the beta rate of about 0.19 new restaurants per restaurant per quarter to roughly 1,200 a quarter at full scale, and it does not explain why the top is an upper bound due to self-selection and stated preferences.

Proposes tests that could failWrongRight
GPT-6.1 Sol · API

Several gates lack explicit measurement windows or clear actions for every outcome, such as the 25% lift gate and the 10% monthly adoption gate.

Gemini 3.8 Flash · API

It provides numeric gates and kill criteria with windows and actions, though some thresholds are arbitrary and not fully tied to the supplied beta rate.

All got right 4

Addresses the actual decisionRightRight
GPT-6.1 Sol · API

It commits early to a worker-led network as the 1000x bet and states stop conditions that would change the call.

Gemini 3.8 Flash · API

It commits early to the worker-led distribution loop and against the AI operating system, with gates and kill criteria that would change the call.

Finds the mechanism the data hidesRightRight
GPT-6.1 Sol · API

It centres the worker as the distribution dimension and supports it with the 41% second-job share and 57 faster, cheaper worker-led signups.

Gemini 3.8 Flash · API

It correctly centers cross-employer worker virality, citing the 41% second-job share, 57 worker-led signups, 9-day close, and lower sales cost.

Answers the CEO's versionRightRight
GPT-6.1 Sol · API

It respectfully declines the AI operating system roadmap, explains the lack of mechanism and squad cost, and leaves room for AI later on proven workflows.

Gemini 3.8 Flash · API

It gives a clear, respectful not-now answer to the AI operating system, grounding it in lack of mechanism, squad spread, and Crewbook, while allowing later labor-focused AI.

Extreme, then back to buildableRightRight
GPT-6.1 Sol · API

It pushes the extreme version to a portable worker identity and earned-tip payout network, then works back to consented linking, invitation tests, and payout gates.

Gemini 3.8 Flash · API

It pushes the worker dimension to a portable worker identity and instant earned-tip liquidity, then works back to a buildable first quarter with unified worker ID and earned-tip payouts.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.295.52None
2GPT-6.1 SolwithAPI92.786.72None
3GPT-6 LunawithAPI88.278.82None
4Opus 5.5withClaude90.774.62None
5Sonnet 5.5withAPI90.265.92None
6Gemini 3.8 FlashwithAPI71.134.521 capped
7Gemini 3.5 Flash-LitewithGemini61.417.02None

About the task

The PM job

Finding the bigger version of a good idea.

Why it matters

Ambition without mechanism is fan fiction. The useful version pushes to the extreme, then works back to something buildable.

What good looks like

  • Names the mechanism that scales
  • Keeps the core insight
  • Works back to a first step you could build

Deliberately not measured

    Capability tested

    Ambitious expansion

    The failure we’re looking for

    Bigger adjectives, same idea

    Grading

    Decision model and LLM judge, calibrated against a blind PM review

    This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.