Tasks / Discover

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The summary is complete, usable, and would let a PM act with only light edits.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  2. Keeps dissent visible100% pass
    The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  3. Respects explicit constraints95% pass
    It is a findings summary for the product team and is under the 600-word limit.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams

Where it slips

  1. Says how many sources support each finding46% pass
    Findings are presented without source counts (e.g., 'five of eight'), using vague terms like 'consistent reports' instead of sizing each theme.
    Gemini 3.8 Flash · API · Eight calls with finance teams
  2. Avoids unsupported claims50% pass
    The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.
    GPT-6.1 Sol · API · Eight calls with finance teams
  3. Uses the supplied evidence correctly63% pass
    The output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
    Opus 5.5 · Claude · Eight calls with finance teams

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Synthesise the eight discovery calls below with finance leads about month-end close. Write a findings summary for the product team, who are deciding whether to build a reconciliation product: what we learned, and how confident we can be in it. Keep it under 600 words.

What the model was given9 items: Scenario, Call 1: Financial controller, logistics company (180 staff), Call 2: Head of finance, SaaS company (95 staff), Call 3: Finance manager, retail chain (400 staff), Call 4: CFO, manufacturing company (250 staff), Call 5: Accountant, agency group (120 staff), Call 6: Finance director, charity (70 staff), Call 7: Controller, hospitality group (300 staff), Call 8: Financial controller, construction firm (210 staff)
ScenarioWe make spend-management software for mid-sized companies and are exploring a reconciliation product. Our PM ran these 40-minute calls in August. Below are the relevant excerpts from each, lightly edited.
Call 1: Financial controller, logistics company (180 staff)“Close takes us eight working days and at least three of those are reconciliation. The bank feed drops transactions, so we tick and bash against statements in Excel.” Asked what she'd pay to fix it: “If it gave me two days back, it pays for itself. But I've heard that before.”
Call 2: Head of finance, SaaS company (95 staff)“Honestly close is fine. We moved to a proper ERP two years ago, it's five days and nobody's complaining.” Later: “The only annoying bit is chasing people for receipts, not reconciling.”
Call 3: Finance manager, retail chain (400 staff)“Card transactions are the nightmare. Hundreds of store cards, the statements come in as PDFs, someone re-keys them. We found £14k of duplicates last quarter.” She reconciles card spend in a shared spreadsheet with four people editing it.
Call 4: CFO, manufacturing company (250 staff)“We switched reconciliation tools last year and I regret it. The migration took four months, we lost all our matching rules, and half the team quietly went back to spreadsheets.” Asked if he'd switch again: “Not unless someone else does the migration for us. Never again otherwise.”
Call 5: Accountant, agency group (120 staff)“Intercompany is where it breaks. Three entities, money moving between them every week, and the balances never agree at month end. I spend a day and a half on it.”
Call 6: Finance director, charity (70 staff)“Close is fine for us. We're small, two of us, it takes four days and we know every transaction.” When asked about tools: “We'd not pay for anything new this year; budgets are frozen.”
Call 7: Controller, hospitality group (300 staff)“Payment processor payouts are the worst: they net off fees and refunds, so nothing matches the bank line. We reconcile by hand and it's always the last thing holding up close.”
Call 8: Financial controller, construction firm (210 staff)“Reconciliation is about 40% of my close. It's not hard, it's just volume, matching the same things every month.” On automation: “I'd want to see it match our messy data before I believed it. Every demo uses clean data.”
What a strong answer doesThe answer key the graders mark against

Themes with honest counts (five of eight describe reconciliation pain, from different sources), the two who say close is fine and the CFO's switching regret kept visible, and hypotheses about willingness to pay and switching cost labelled as such.

Critical failures (cap the score at 40)
  • Invents a quote
Case

Mid-level · v1.9 · anonymised real · B2B, finance

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit7 of 9 checks passed

The only error is a stray year in the title; the rest of the reasoning, evidence, and structure is solid and usable immediately after deleting '2026'.

Output

Reconciliation discovery: findings from 8 finance-lead calls (August 2026)3

Bottom line: Most of the finance leads we spoke to have real reconciliation pain. But it shows up as several different problems, not one, and the biggest barrier to adoption is distrust and switching cost. We have almost no evidence on willingness to pay. This justifies more targeted discovery. It doesn't yet justify a decision to build.

What we learned

  1. The pain is common but not universal. Six of eight say reconciliation is a significant part of close (Calls 1, 3, 4, 5, 7, 8).1 The two who said close is fine (2, 6) run the smallest companies, at 95 and 70 staff, and one of them has a modern ERP. That hints at a size threshold around 100+ staff, but it rests on two data points.

2. Where people gave numbers, the cost is large. - About 3 of 8 close days (1) - 1.5 days a month (5) - About 40% of close (8) - £14k of duplicates found in one quarter (3)

All of these are self-reported.

3. "Reconciliation" is at least five separate problems. Each call with pain named a different source: - Bank feed dropping transactions (1) - Card statements arriving as PDFs and re-keyed into a shared spreadsheet (3) - Intercompany balances that don't agree across entities (5) - Payment-processor payouts that net off fees and refunds, so nothing matches the bank line (7) - Sheer volume of repetitive matching (8)

No two calls named the same source. One "reconciliation product" would have to handle different matching logic for different segments. Card spend (3) sits closest to what we already do and could be a natural starting point, but it came up in only one call.

4. Switching cost and skepticism are the main barrier.2 - Call 4 switched reconciliation tools last year. Migration took four months, they lost their matching rules, and half the team drifted back to spreadsheets. He won't switch again unless someone else does the migration. - Call 8 won't believe a tool works until it handles their messy data, because demos always use clean data. - Call 1 has "heard that before."

That's three of eight, including people with heavy pain. Done-for-you migration and proof on the customer's own data look like requirements, not differentiators.

  1. One adjacent signal. Call 2's only complaint was chasing people for receipts. That's our existing spend-management territory, not reconciliation.

How confident can we be?

Moderate confidence: - Reconciliation takes a meaningful share of close time for mid-sized finance teams. - The underlying causes vary widely between companies. - Buyers are wary of switching tools and of vendor demos.

Low or no confidence: - Willingness to pay. Only one person answered the question, and only conditionally. One has a budget freeze. - Which problem to build for. Each pain source appears exactly once. - Size threshold. It rests on two calls. - Prevalence. We don't know how these eight were recruited, whether they're customers, or how representative they are.

Method limitations: - Eight calls, one interviewer, one month. - We're working from excerpts chosen as "relevant," so the full transcripts may shift the emphasis. - The pay question was open-ended and hypothetical.

What would raise confidence before a build decision

  • Recruit 8–10 more finance leads at companies with 150–400 staff. Screen them by pain source so we learn which problem repeats.
  • Test the card-spend starting point with existing customers. Ask how they reconcile card spend today. We may already hold the data needed to solve it.
  • Test willingness to pay concretely. Use price anchors or offer a paid pilot rather than asking an open question.
  • Run a matching test on one prospect's real, messy data. Call 8 set that bar, and it's the test that would convince the skeptics.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Call 4: CFO, manufacturing company (250 staff)We switched reconciliation tools last year and I regret it.

    What we’d changeRecount without Call 4, which is about a failed tool switch, not close time. It's five of eight.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent it as a concern three people raised. Eight calls can't rank the barriers.

  3. 3Constraint missedRestore the constraintQuick edit

    Source: BriefKeep it under 600 words.

    What we’d changeTrim it: it runs past the 600-word limit. And drop the year, which the brief doesn't give.

Claims the judge couldn’t find in the brief

  • The discovery calls took place in August 2026.

Grades and run details

Decision model 72 · LLM judge 8 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly22%
  • passAddresses the actual decision97%
  • partialRespects explicit constraints20%
  • passIdentifies material uncertainty99%
  • failAvoids unsupported claims18%
  • passProduces the required deliverable50%
  • passKeeps dissent visible51%
  • passWeights behaviour over opinion88%
  • passSays how many sources support each finding36%
Run
Run
#1
Time to output
35 s
Submitted
24 Sept 2026

Usable with a quick edit8 of 9 checks passed

Only a wording edit is needed to hedge the adoption forecast; the evidence and reasoning are sound and would not need re-checking.

Output

Findings: reconciliation discovery

Eight 40-minute calls with finance leads, August

Bottom line

These calls support further validation of a targeted reconciliation product, not yet a broad build decision. Several participants described substantial manual work, but “reconciliation” covers different problems. Adoption will depend on handling messy data and reducing switching effort—not just offering automated matching.

What we learned

1. Reconciliation is a material problem for some, not all. Five participants described current reconciliation pain: - Logistics: three or more days of an eight-day close, with missing bank-feed transactions checked against statements in Excel. - Retail: manual re-keying of card-statement PDFs and a shared spreadsheet used by four people; £14k of duplicates found last quarter. - Agency: intercompany balances take a day and a half to resolve. - Hospitality: processor payouts net of fees and refunds require manual reconciliation and hold up close. - Construction: reconciliation accounts for roughly 40% of close, driven by repetitive matching volume.

Two participants—SaaS and charity—said close was fine. The SaaS participant’s remaining frustration was receipt chasing, not reconciliation. The manufacturing participant primarily described a failed tool migration rather than quantifying current close pain.

2. There is no single, validated product scope yet. The examples span incomplete bank feeds, card-statement ingestion, intercompany balances, processor payouts and high-volume matching. These may require different integrations and workflows. Card reconciliation is plausibly adjacent to our spend-management product, but its evidence here comes from one call; we should not assume it represents the wider opportunity.

3. Trust and implementation are central adoption requirements. The construction controller wanted proof using their own messy data. The logistics controller expressed scepticism about promised time savings. The manufacturing CFO reported a four-month migration, lost matching rules and a return to spreadsheets; he would not switch again without migration support.

Together, these suggest that realistic-data trials, continuity of existing rules and migration assistance may matter as much as matching capability.

4. Willingness to pay remains unproven. One participant said saving two days would make a solution worthwhile, but gave no price or commitment. Another explicitly ruled out new spending this year. The duplicate-loss example indicates potential financial value, but does not establish recoverable savings or a purchasing budget.

How confident can we be?

  • Moderate confidence that manual reconciliation creates meaningful pain for a subset of companies like these, and that trust and switching friction deserve early attention.
  • Low confidence in market prevalence, the best initial segment, willingness to pay, achievable automation rates or expected close-time savings.

This is a small qualitative sample; recruitment and representativeness are unspecified. Evidence is self-reported, not verified through workflow observation or transaction data. The calls establish neither demand for one common solution nor demand specifically from our existing customers.

Recommended next step

Choose a narrow use case to test—not a general reconciliation platform. Observe workflows and trial matching on real, messy data with several similar customers. Measure manual effort removed, exceptions remaining and implementation burden, then test a priced pilot. Advance to a build decision only with repeatable evidence of value and credible adoption commitments.

Grades and run details

Decision model 94 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly80%
  • passAddresses the actual decision98%
  • passRespects explicit constraints47%
  • passIdentifies material uncertainty99%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable79%
  • passKeeps dissent visible96%
  • passWeights behaviour over opinion91%
  • passSays how many sources support each finding71%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 9 checks passed

The output is accurate, well-structured, and actionable with only minor wording changes needed.

Output

Findings: month-end reconciliation

These calls justify testing a focused reconciliation proposition, but do not yet justify building a broad product. Six participants described reconciliation pain or dissatisfaction with their tools. However, their problems differ substantially, and evidence of willingness to adopt or pay is weak.

What we learned

Reconciliation can materially delay close, but it is not a universal problem. The logistics controller reported at least three reconciliation days within an eight-day close. The agency accountant spends 1.5 days on intercompany balances. The construction controller estimated reconciliation at 40% of close, while the hospitality controller said processor payouts are the final blocker.

Conversely, the SaaS and charity participants explicitly said close works well. SaaS’s remaining frustration was receipt chasing, which should not be counted as reconciliation demand. These counterexamples argue against targeting all finance teams simply because they run a monthly close.

“Reconciliation” covers several distinct jobs.

ProblemEvidence from calls
Incomplete transaction dataLogistics: missing bank-feed transactions require checking statements in Excel
Card-statement processingRetail: PDF re-keying, shared spreadsheet coordination and £14k of duplicates found last quarter
Intercompany balancingAgency: transfers across three entities create mismatched balances
Processor payout reconciliationHospitality: fees and refunds obscure the relationship between payments and bank deposits
Repetitive matching at volumeConstruction: recurring manual work, described as time-consuming rather than difficult

These may require different data integrations, matching logic and workflows. Combining them into one demand signal would overstate support for any particular product.

Spreadsheets persist for different reasons. Logistics uses them to compensate for feed gaps; retail uses them to coordinate work; manufacturing returned to them after a failed tool migration. Spreadsheet use alone does not establish a desire to replace spreadsheets.

Trust and migration could determine adoption. The construction controller wants proof using messy, real data. The logistics controller is sceptical of promised savings. The manufacturing CFO made migration assistance a condition of switching after losing matching rules during a four-month migration. These are specific adoption requirements, although we cannot assume every buyer shares them.

How confident can we be?

Moderate confidence that meaningful problems exist in some of these organisations. Participants supplied concrete workflows, time estimates and one financial example. These remain self-reports: we have not observed the work or verified the figures. The £14k in duplicates does not establish losses incurred or savings our product could deliver.

Low confidence about prevalence, the best segment or a common solution. Eight exploratory calls provide useful hypotheses, not market estimates. Recruitment criteria and customer relationships are unspecified. Each industry appears only once, so differences cannot reliably be attributed to sector, company size or ERP maturity.

Very low confidence about commercial demand. One participant gave a conditional value statement after being asked about price, without naming a budget or committing to buy. Another explicitly ruled out new spending this year. No purchase, pilot commitment or switching decision is evidenced.

Implication for the product decision

Prioritise validation of card-spend reconciliation, given its proximity to our existing spend-management product, while recognising that only one call directly supports this use case.

Recruit more teams with that workflow, inspect actual statements and reconciliation records, and test a narrow prototype on messy data. Measure time saved, incorrect matches and exception-handling effort. Seek concrete pilot commitments and establish budget ownership and migration requirements before committing to build.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly46%
  • passAddresses the actual decision95%
  • passRespects explicit constraints41%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable71%
  • passKeeps dissent visible86%
  • passWeights behaviour over opinion87%
  • passSays how many sources support each finding53%
Run
Run
#1
Time to output
35 s
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Uses the supplied evidence correctlyWrongRightRight
Opus 5.5 · Claude

The output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.

GPT-6.1 Sol · API

All factual claims and figures come from the supplied call excerpts; no invented current-state facts.

GPT-6 Astra · ChatGPT

All facts and figures used are taken correctly from the supplied call excerpts with no invented current-state details.

Avoids unsupported claimsWrongWrongRight
Opus 5.5 · Claude

The date 'August 2026' is presented as a confident fact without any basis in the supplied context.

GPT-6.1 Sol · API

The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.

GPT-6 Astra · ChatGPT

Interpretations and hypotheses are hedged or labelled as such, and confident claims are limited to those supported by the excerpts.

All got right 7

Addresses the actual decisionRightRightRight
Opus 5.5 · Claude

The output clearly recommends that the evidence doesn't yet justify building and that more targeted discovery is needed, and states what would change that call.

GPT-6.1 Sol · API

Commits to further validation of a targeted product rather than a broad build, and specifies repeatable evidence/adoption commitments as the condition to change.

GPT-6 Astra · ChatGPT

The output clearly commits to testing a focused reconciliation proposition rather than building broadly, and states what validation results would change that call.

Respects explicit constraintsRightRightRight
Opus 5.5 · Claude

The output is a findings summary under 600 words for the product team, as requested.

GPT-6.1 Sol · API

Under 600 words, written as a findings summary for the product team, and respects the requested reader and form.

GPT-6 Astra · ChatGPT

It is a findings summary for the product team and is under the 600-word limit.

Identifies material uncertaintyRightRightRight
Opus 5.5 · Claude

It names unknown willingness to pay, problem to build for, size threshold, and recruitment bias, and how to resolve them.

GPT-6.1 Sol · API

Names specific unknowns—market prevalence, segment, willingness to pay, automation rates, savings—and says messy-data trials and priced pilots would resolve them.

GPT-6 Astra · ChatGPT

It names prevalence, segment, willingness to pay, migration needs, and unverified self-reports as open questions and says how pilots and messy-data testing would resolve them.

Produces the required deliverableRightRightRight
Opus 5.5 · Claude

The deliverable is a complete findings summary that a product team could act on with minimal edits.

GPT-6.1 Sol · API

Complete, actionable findings summary within length with a clear next step for the product team.

GPT-6 Astra · ChatGPT

The summary is complete, usable, and would let a PM act with only light edits.

Keeps dissent visibleRightRightRight
Opus 5.5 · Claude

Calls 2 and 6 who said close is fine are explicitly mentioned and kept in view.

GPT-6.1 Sol · API

SaaS and charity close-fine responses and manufacturing regret are explicitly kept visible.

GPT-6 Astra · ChatGPT

The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.

Weights behaviour over opinionRightRightRight
Opus 5.5 · Claude

The output flags self-reported numbers and hypothetical pay answers, and weights the actual migration experience (an observed behaviour) heavily.

GPT-6.1 Sol · API

Distinguishes observed workarounds and migration behaviour from stated intent, and flags willingness to pay as unproven.

GPT-6 Astra · ChatGPT

It separates observed workarounds and past behaviour from stated intent and flags commercial demand as very low confidence.

Says how many sources support each findingRightRightRight
Opus 5.5 · Claude

Every finding is sized by number of calls (e.g., 'six of eight', 'three of eight', 'only one call').

GPT-6.1 Sol · API

Substantive themes include five of eight, two of eight, and named single-call sources; no universal customer consensus is claimed.

GPT-6 Astra · ChatGPT

Findings are anchored to named calls and explicit counts such as 'six participants', 'eight exploratory calls', 'only one call', and 'one participant/another', avoiding unsupported generalisation.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 83% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.495.02None
2GPT-6.1 SolwithAPI94.490.02None
3GPT-6 LunawithAPI91.780.02None
4Sonnet 5.5withAPI94.465.02None
5Opus 5.5withClaude69.470.02None
6Gemini 3.8 FlashwithAPI52.855.02None
7Gemini 3.5 Flash-LitewithGemini58.330.02None

About the task

The PM job

Turning a stack of call transcripts into what we actually learned.

Why it matters

Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.

What good looks like

  • Quotes evidence for each theme and counts sources honestly
  • Weights what customers did above what they say they'd do
  • Keeps important dissent visible
  • Labels hypotheses as hypotheses
  • Says what the research cannot tell us

Deliberately not measured

  • Transcript clean-up
  • Persona illustration
Capability tested

Faithful synthesis of qualitative research

The failure we’re looking for

Invents customer consensus or loses important dissent

Grading

Decision model and LLM judge, calibrated against a blind PM review