Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.0 · 2 casesDifficulty

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model, LLM judge and blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.

Results

Every setup we’ve tested on this task, across all cases and repeats.

#Model · HarnessTask scoreDecision modelLLM judgePM reviewRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.

The brief

We plan to sell an AI sales-development rep, an agent that finds prospects and sends personalised outbound email, to agencies with 5–20 staff. Before we commit a year to it, find the strongest reason this fails. Write it up for the founding team: the one reason, the evidence for it, and what we would need to see to be proved wrong. Keep it under 400 words.

ScenarioWe are a three-person founding team with £800k of pre-seed funding. Planned price: $600 a month per agency.
EvidenceFrom 14 agency interviews: 11 said most new business comes from referrals and repeat clients (on average 80% of new revenue). 9 had tried outbound (cold email or LinkedIn) in the last two years, and 7 of those stopped within six months. Average deal size is $18k with a 6–10 week sales cycle. Two competitors each raised over $10m in the last year, selling mostly to software companies with 50+ staff.
What a strong answer does

The load-bearing assumption is that agencies' growth is outbound-constrained; the referral evidence says it is not.

Case

v1.2 · synthetic · B2B, go-to-market