Challenge an idea
Can the model find the strongest reason an idea may fail, backed by evidence?
The PM job
Pressure-testing a proposal before committing a team to it.
Why it matters
The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.
What good looks like
- Identifies the load-bearing assumption
- Uses the supplied evidence, not generic risks
- Proposes the cheapest way to test the assumption
Deliberately not measured
- Tone
- Number of objections raised
Evidence-based critique
Theatrical negativity without evidence
Decision model, LLM judge and blind PM review
Vanilla prompt (core) · With Roast Me skill
This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
We plan to sell an AI sales-development rep, an agent that finds prospects and sends personalised outbound email, to agencies with 5–20 staff. Before we commit a year to it, find the strongest reason this fails. Write it up for the founding team: the one reason, the evidence for it, and what we would need to see to be proved wrong. Keep it under 400 words.
The load-bearing assumption is that agencies' growth is outbound-constrained; the referral evidence says it is not.
v1.2 · synthetic · B2B, go-to-market