Tasks / Discover

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The summary is complete, usable, and would let a PM act with only light edits.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  2. Keeps dissent visible100% pass
    The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  3. Respects explicit constraints95% pass
    It is a findings summary for the product team and is under the 600-word limit.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams

Where it slips

  1. Says how many sources support each finding46% pass
    Findings are presented without source counts (e.g., 'five of eight'), using vague terms like 'consistent reports' instead of sizing each theme.
    Gemini 3.8 Flash · API · Eight calls with finance teams
  2. Avoids unsupported claims50% pass
    The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.
    GPT-6.1 Sol · API · Eight calls with finance teams
  3. Uses the supplied evidence correctly63% pass
    The output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
    Opus 5.5 · Claude · Eight calls with finance teams

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Synthesise the eight discovery calls below with finance leads about month-end close. Write a findings summary for the product team, who are deciding whether to build a reconciliation product: what we learned, and how confident we can be in it. Keep it under 600 words.

What the model was given9 items: Scenario, Call 1: Financial controller, logistics company (180 staff), Call 2: Head of finance, SaaS company (95 staff), Call 3: Finance manager, retail chain (400 staff), Call 4: CFO, manufacturing company (250 staff), Call 5: Accountant, agency group (120 staff), Call 6: Finance director, charity (70 staff), Call 7: Controller, hospitality group (300 staff), Call 8: Financial controller, construction firm (210 staff)
ScenarioWe make spend-management software for mid-sized companies and are exploring a reconciliation product. Our PM ran these 40-minute calls in August. Below are the relevant excerpts from each, lightly edited.
Call 1: Financial controller, logistics company (180 staff)“Close takes us eight working days and at least three of those are reconciliation. The bank feed drops transactions, so we tick and bash against statements in Excel.” Asked what she'd pay to fix it: “If it gave me two days back, it pays for itself. But I've heard that before.”
Call 2: Head of finance, SaaS company (95 staff)“Honestly close is fine. We moved to a proper ERP two years ago, it's five days and nobody's complaining.” Later: “The only annoying bit is chasing people for receipts, not reconciling.”
Call 3: Finance manager, retail chain (400 staff)“Card transactions are the nightmare. Hundreds of store cards, the statements come in as PDFs, someone re-keys them. We found £14k of duplicates last quarter.” She reconciles card spend in a shared spreadsheet with four people editing it.
Call 4: CFO, manufacturing company (250 staff)“We switched reconciliation tools last year and I regret it. The migration took four months, we lost all our matching rules, and half the team quietly went back to spreadsheets.” Asked if he'd switch again: “Not unless someone else does the migration for us. Never again otherwise.”
Call 5: Accountant, agency group (120 staff)“Intercompany is where it breaks. Three entities, money moving between them every week, and the balances never agree at month end. I spend a day and a half on it.”
Call 6: Finance director, charity (70 staff)“Close is fine for us. We're small, two of us, it takes four days and we know every transaction.” When asked about tools: “We'd not pay for anything new this year; budgets are frozen.”
Call 7: Controller, hospitality group (300 staff)“Payment processor payouts are the worst: they net off fees and refunds, so nothing matches the bank line. We reconcile by hand and it's always the last thing holding up close.”
Call 8: Financial controller, construction firm (210 staff)“Reconciliation is about 40% of my close. It's not hard, it's just volume, matching the same things every month.” On automation: “I'd want to see it match our messy data before I believed it. Every demo uses clean data.”
What a strong answer doesThe answer key the graders mark against

Themes with honest counts (five of eight describe reconciliation pain, from different sources), the two who say close is fine and the CFO's switching regret kept visible, and hypotheses about willingness to pay and switching cost labelled as such.

Critical failures (cap the score at 40)
  • Invents a quote
Case

Mid-level · v1.9 · anonymised real · B2B, finance

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit7 of 9 checks passed

The output lacks required source counts per theme, so a PM would need to re-check the evidence to add them, which is more than a quick edit.

Output

Executive Summary

We interviewed eight finance leads across various industries (70–400 FTEs) to assess whether to build a reconciliation product.

The takeaway: A generic, standalone reconciliation tool is a high-risk bet with massive adoption hurdles. However, there is acute, high-value pain around messy, high-volume card and payment processor reconciliations—areas adjacent to our core spend-management platform.

---

What We Learned

#### 1. Pain is driven by transaction complexity, not company size * High-pain profiles (FTEs 180–400): Companies with complex transaction streams (retail cards, payment gateway net-offs, intercompany transfers, dropped bank feeds) spend 30–40% of their month-end close (1.5 to 3+ days) manually ticking spreadsheets. * Low-pain profiles (FTEs <120): Smaller teams or digitally native businesses with modern ERPs report manageable 4-to-5-day closes. Their bottlenecks lie in chasing receipts, not reconciling balances.

#### 2. Buyers are deeply skeptical and burned by past tools * Migration trauma: One CFO spent four months migrating to a dedicated reconciliation tool, lost matching rules, and saw the team revert to Excel. He will not switch again without full migration support. * “Clean data” cynicism: Buyers repeatedly noted that vendors promise automated matching but fail on real-world messy data (e.g., processor fees netted against payouts, missing bank feed lines).

#### 3. Clear synergy with core spend management Retail and hospitality leads cited corporate cards and payment payouts as their primary close bottlenecks (e.g., one retail chain discovered £14k in duplicate card transactions via shared spreadsheets). Reconciling card spend naturally fits our existing product footprint.

---

Confidence Assessment

AreaConfidence LevelRationale
:---:---:---
Pain existenceHighConsistent reports of spreadsheet-based workarounds consuming days of close time in transaction-heavy businesses.
Strategic alignmentHighCard transaction and payout reconciliation aligns directly with our spend-management domain.
Willingness to buy a standalone toolLowHigh switching costs, migration friction, and tool fatigue make buyers reluctant to purchase an unproven standalone system.
Willingness to payMediumBuyers indicate solutions “pay for themselves” if they save 2+ days, but budgets at the lower end (e.g., non-profits) are frozen.

Sample Limitations: Eight calls provide qualitative directional signals, but the sample spans very diverse business models (logistics, SaaS, retail, charity, construction). We lack statistical significance on pricing thresholds and specific ERP integrations.

---

Product Recommendation

  1. Do not build a generic, horizontal reconciliation platform. Competing against established ERPs and broad reconciliation players carries extreme migration friction.
  2. Double down on native spend reconciliation: Scope a feature set within our existing platform focused on automated card matching and statement ingestion (e.g., parsing PDF card statements, catching duplicate charges).
  3. De-risk with a messy-data PoC: Before committing to a major build, run a technical pilot with 2–3 design partners (such as the retail chain or logistics firm) to prove our matching algorithm handles their dirtiest statement data without manual rule configuration.

Grades and run details

Decision model 61 · LLM judge 8 of 10 checks
Decision model checks
  • partialUses the supplied evidence correctly3%
  • passAddresses the actual decision89%
  • partialRespects explicit constraints39%
  • passIdentifies material uncertainty37%
  • failAvoids unsupported claims55%
  • passProduces the required deliverable59%
  • passKeeps dissent visible25%
  • partialWeights behaviour over opinion34%
  • failSays how many sources support each finding39%
Run
Run
#1
API response time
13 s
Submitted
1 Oct 2026

Usable with a quick edit8 of 9 checks passed

Only a wording edit is needed to hedge the adoption forecast; the evidence and reasoning are sound and would not need re-checking.

Output

Findings: reconciliation discovery

Eight 40-minute calls with finance leads, August

Bottom line

These calls support further validation of a targeted reconciliation product, not yet a broad build decision. Several participants described substantial manual work, but “reconciliation” covers different problems. Adoption will depend on handling messy data and reducing switching effort—not just offering automated matching.

What we learned

1. Reconciliation is a material problem for some, not all. Five participants described current reconciliation pain: - Logistics: three or more days of an eight-day close, with missing bank-feed transactions checked against statements in Excel. - Retail: manual re-keying of card-statement PDFs and a shared spreadsheet used by four people; £14k of duplicates found last quarter. - Agency: intercompany balances take a day and a half to resolve. - Hospitality: processor payouts net of fees and refunds require manual reconciliation and hold up close. - Construction: reconciliation accounts for roughly 40% of close, driven by repetitive matching volume.

Two participants—SaaS and charity—said close was fine. The SaaS participant’s remaining frustration was receipt chasing, not reconciliation. The manufacturing participant primarily described a failed tool migration rather than quantifying current close pain.

2. There is no single, validated product scope yet. The examples span incomplete bank feeds, card-statement ingestion, intercompany balances, processor payouts and high-volume matching. These may require different integrations and workflows. Card reconciliation is plausibly adjacent to our spend-management product, but its evidence here comes from one call; we should not assume it represents the wider opportunity.

3. Trust and implementation are central adoption requirements. The construction controller wanted proof using their own messy data. The logistics controller expressed scepticism about promised time savings. The manufacturing CFO reported a four-month migration, lost matching rules and a return to spreadsheets; he would not switch again without migration support.

Together, these suggest that realistic-data trials, continuity of existing rules and migration assistance may matter as much as matching capability.

4. Willingness to pay remains unproven. One participant said saving two days would make a solution worthwhile, but gave no price or commitment. Another explicitly ruled out new spending this year. The duplicate-loss example indicates potential financial value, but does not establish recoverable savings or a purchasing budget.

How confident can we be?

  • Moderate confidence that manual reconciliation creates meaningful pain for a subset of companies like these, and that trust and switching friction deserve early attention.
  • Low confidence in market prevalence, the best initial segment, willingness to pay, achievable automation rates or expected close-time savings.

This is a small qualitative sample; recruitment and representativeness are unspecified. Evidence is self-reported, not verified through workflow observation or transaction data. The calls establish neither demand for one common solution nor demand specifically from our existing customers.

Recommended next step

Choose a narrow use case to test—not a general reconciliation platform. Observe workflows and trial matching on real, messy data with several similar customers. Measure manual effort removed, exceptions remaining and implementation burden, then test a priced pilot. Advance to a build decision only with repeatable evidence of value and credible adoption commitments.

Grades and run details

Decision model 94 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly80%
  • passAddresses the actual decision98%
  • passRespects explicit constraints47%
  • passIdentifies material uncertainty99%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable79%
  • passKeeps dissent visible96%
  • passWeights behaviour over opinion91%
  • passSays how many sources support each finding71%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 9 checks passed

The output is evidence-backed and actionable as-is. A PM would only need to add an explicit build trigger, which is a light edit.

Output

Month-end close discovery: findings summary

Bottom line

Reconciliation is a real, recurring pain for mid-sized finance teams, but these eight calls don't yet show which product to build or whether anyone will pay for it. I'd treat this as grounds for a further, more targeted round of research, not a build decision.

What we learned

1. Five of eight leads have a real reconciliation problem (calls 1, 3, 5, 7, 8). Two said close is fine (calls 2 and 6), and one has a tool and regrets it (call 4).

2. They describe different problems, not one. - Bank feeds dropping transactions, matched by hand in Excel (call 1) - Card statements arriving as PDFs and re-keyed, with £14k of duplicates in a quarter (call 3) - Intercompany balances that never agree, costing about 1.5 days (call 5) - Payment processor payouts that net off fees and refunds, and are always the last thing holding up close (call 7) - Sheer matching volume, about 40% of one controller's close (call 8)

No problem appears in more than one call, so we can't yet say which is most common or most valuable.

3. The time cost is meaningful but self-reported. Estimates ranged from 1.5 days to 3 of 8 days. These are recollections, not measurements.

4. Those who are fine tend to be smaller or better equipped. The two "close is fine" leads are the smallest companies in the sample (70 and 95 staff), and one already has a modern ERP. This is suggestive at best with eight calls. The pain sits mostly in the 180–400 staff range.

5. Adoption barriers are as prominent as the pain. - Call 4's failed migration (four months, lost matching rules, team back on spreadsheets) is a warning about switching costs. He would only switch again if someone else did the migration. - Call 8 won't believe automation works until he sees it handle messy data: "Every demo uses clean data." - Call 1 is sceptical: "I've heard that before."

6. Willingness to pay is essentially untested. Only call 1 gave a conditional answer ("if it gave me two days back, it pays for itself"), and she hedged it. Call 6 has a budget freeze. No one expressed urgency or asked about pricing.

7. An adjacent finding: call 2's only complaint was chasing receipts, which sits closer to our existing spend-management product than reconciliation does.

How confident can we be?

ClaimConfidence
Reconciliation is a meaningful time sink for some mid-sized teamsModerate
Pain varies by reconciliation typeModerate to high (consistent across calls)
Size of the time savingLow (self-reported, varied definitions)
Willingness to payLow (one hedged signal)
Which segment or use case to targetLow
Smaller or ERP-equipped firms are a weaker fitLow (three data points)

Other limitations: - Eight calls is a small sample, and we don't know how participants were recruited, so existing customers or those already interested may be overrepresented. - The excerpts are edited and one interviewer ran all the calls. - Hypothetical questions ("would you pay?") overstate real intent.

Suggested next steps

  1. Pick a wedge. Card reconciliation (call 3) is closest to our current product. Follow up with calls 3, 5 and 7 to size how often each problem occurs and what it costs.
  2. Test with real data. Ask call 8 and similar leads for anonymised, messy samples to see whether matching is feasible.
  3. Design for migration. Call 4 suggests onboarding and preserving matching rules may decide adoption. Treat this as a product requirement, not an afterthought.
  4. Test price and urgency. Use concrete pricing or a paid pilot, not hypotheticals.
  5. Widen the sample. Include firms of 150–500 staff and a mix of ERP and non-ERP setups, and record how each participant was recruited.

Grades and run details

Decision model 100 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly59%
  • passAddresses the actual decision93%
  • passRespects explicit constraints43%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims31%
  • passProduces the required deliverable78%
  • passKeeps dissent visible98%
  • passWeights behaviour over opinion98%
  • passSays how many sources support each finding65%
Run
Run
#1
API response time
18 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Addresses the actual decisionRightRightMixed
Gemini 3.8 Flash · API

Commits unambiguously to not building a generic standalone tool and instead building native spend reconciliation with a pilot, and states the pilot condition that would change the call.

GPT-6.1 Sol · API

Commits to further validation of a targeted product rather than a broad build, and specifies repeatable evidence/adoption commitments as the condition to change.

Sonnet 5.5 · API

It recommends more research but does not state what specific result or condition would trigger a build decision.

Avoids unsupported claimsMixedWrongRight
Gemini 3.8 Flash · API

Interpretations are presented as findings or confidence assessments, not as established facts, and are grounded in the supplied evidence.

GPT-6.1 Sol · API

The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.

Sonnet 5.5 · API

Hypotheses and limitations are labelled as suggestive or untested, not as established fact.

Says how many sources support each findingWrongRightRight
Gemini 3.8 Flash · API

Findings are presented without source counts (e.g., 'five of eight'), using vague terms like 'consistent reports' instead of sizing each theme.

GPT-6.1 Sol · API

Substantive themes include five of eight, two of eight, and named single-call sources; no universal customer consensus is claimed.

Sonnet 5.5 · API

Findings consistently include source counts ('five of eight', 'two said', 'call 4').

All got right 6

Uses the supplied evidence correctlyRightRightRight
Gemini 3.8 Flash · API

All factual claims about the current situation are taken directly from the brief or derived by arithmetic from it, with no inventions.

GPT-6.1 Sol · API

All factual claims and figures come from the supplied call excerpts; no invented current-state facts.

Sonnet 5.5 · API

All factual statements trace to the supplied calls, with no invented current-situation facts.

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

Output is a findings summary under 600 words, addressed to the product team, respecting the requested form and length.

GPT-6.1 Sol · API

Under 600 words, written as a findings summary for the product team, and respects the requested reader and form.

Sonnet 5.5 · API

The output is a findings summary for the product team and is under the 600-word limit.

Identifies material uncertaintyRightRightRight
Gemini 3.8 Flash · API

Names sample limitations, lack of statistical significance on pricing and ERP integrations, and proposes a messy-data pilot to resolve the key uncertainty.

GPT-6.1 Sol · API

Names specific unknowns—market prevalence, segment, willingness to pay, automation rates, savings—and says messy-data trials and priced pilots would resolve them.

Sonnet 5.5 · API

It names unknown WTP, segment, and time savings, and says how to resolve them in next steps.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

Complete findings summary with confidence assessment and recommendation, under 600 words, directly usable by the product team.

GPT-6.1 Sol · API

Complete, actionable findings summary within length with a clear next step for the product team.

Sonnet 5.5 · API

The summary is complete, readable, and directly usable by the product team.

Keeps dissent visibleRightRightRight
Gemini 3.8 Flash · API

The two low-pain profiles (calls 2 and 6) and the CFO's switching regret (call 4) are explicitly included and not buried.

GPT-6.1 Sol · API

SaaS and charity close-fine responses and manufacturing regret are explicitly kept visible.

Sonnet 5.5 · API

The two close-is-fine calls and the CFO's switching regret are kept visible throughout.

Weights behaviour over opinionRightRightRight
Gemini 3.8 Flash · API

Weighs observed behavior (spreadsheet workarounds, migration reversion) over stated intent, and flags willingness-to-pay as medium confidence based on stated preferences.

GPT-6.1 Sol · API

Distinguishes observed workarounds and migration behaviour from stated intent, and flags willingness to pay as unproven.

Sonnet 5.5 · API

It flags self-reported recollections and hypothetical pricing as weaker than observed behaviours and workarounds.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 83% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.495.02None
2GPT-6.1 SolwithAPI94.490.02None
3GPT-6 LunawithAPI91.780.02None
4Sonnet 5.5withAPI94.465.02None
5Opus 5.5withClaude69.470.02None
6Gemini 3.8 FlashwithAPI52.855.02None
7Gemini 3.5 Flash-LitewithGemini58.330.02None

About the task

The PM job

Turning a stack of call transcripts into what we actually learned.

Why it matters

Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.

What good looks like

  • Quotes evidence for each theme and counts sources honestly
  • Weights what customers did above what they say they'd do
  • Keeps important dissent visible
  • Labels hypotheses as hypotheses
  • Says what the research cannot tell us

Deliberately not measured

  • Transcript clean-up
  • Persona illustration
Capability tested

Faithful synthesis of qualitative research

The failure we’re looking for

Invents customer consensus or loses important dissent

Grading

Decision model and LLM judge, calibrated against a blind PM review