Extract discovery insights
Can the model separate evidence, themes and hypotheses without inventing consensus?
The PM job
Turning a stack of call transcripts into what we actually learned.
Why it matters
Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.
What good looks like
- Quotes evidence for each theme and counts sources honestly
- Keeps important dissent visible
- Labels hypotheses as hypotheses
- Says what the research cannot tell us
Deliberately not measured
- Transcript clean-up
- Persona illustration
Faithful synthesis of qualitative research
Invents customer consensus or loses important dissent
Decision model, LLM judge and blind PM review
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
Synthesise the eight discovery calls below with finance leads about month-end close. Write a findings summary for the product team, who are deciding whether to build a reconciliation product: what we learned, and how confident we can be in it. Keep it under 600 words.
Themes with honest counts (five of eight describe reconciliation pain, from different sources), the two who say close is fine and the CFO's switching regret kept visible, and hypotheses about willingness to pay and switching cost labelled as such.
- Invents a quote
v1.2 · anonymised real · B2B, finance