Challenge an idea
Can the model find the strongest reason an idea may fail, backed by evidence?
The PM job
Pressure-testing a proposal before committing a team to it.
Why it matters
The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.
What good looks like
- Identifies the load-bearing assumption
- Uses the supplied evidence, not generic risks
- Proposes the cheapest way to test the assumption
Deliberately not measured
- Tone
- Number of objections raised
Evidence-based critique
Theatrical negativity without evidence
Decision model, LLM judge and blind PM review
Vanilla prompt (core) · With Roast Me skill
This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
Our product team proposes adding a community forum to our budgeting app, described below. Challenge the proposal: write the strongest case against building it as proposed, for the product lead. Keep it under 400 words.
The privacy norm undermines participation, so the forum will not produce the retention it assumes; the real, evidenced need is self-serve help content for the top 'how do I' topics.
v1.2 · synthetic · consumer, fintech