Tasks / Discover

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The summary is complete, usable, and would let a PM act with only light edits.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  2. Keeps dissent visible100% pass
    The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  3. Respects explicit constraints95% pass
    It is a findings summary for the product team and is under the 600-word limit.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams

Where it slips

  1. Says how many sources support each finding46% pass
    Findings are presented without source counts (e.g., 'five of eight'), using vague terms like 'consistent reports' instead of sizing each theme.
    Gemini 3.8 Flash · API · Eight calls with finance teams
  2. Avoids unsupported claims50% pass
    The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.
    GPT-6.1 Sol · API · Eight calls with finance teams
  3. Uses the supplied evidence correctly63% pass
    The output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
    Opus 5.5 · Claude · Eight calls with finance teams

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Synthesise the eight discovery calls below with finance leads about month-end close. Write a findings summary for the product team, who are deciding whether to build a reconciliation product: what we learned, and how confident we can be in it. Keep it under 600 words.

What the model was given9 items: Scenario, Call 1: Financial controller, logistics company (180 staff), Call 2: Head of finance, SaaS company (95 staff), Call 3: Finance manager, retail chain (400 staff), Call 4: CFO, manufacturing company (250 staff), Call 5: Accountant, agency group (120 staff), Call 6: Finance director, charity (70 staff), Call 7: Controller, hospitality group (300 staff), Call 8: Financial controller, construction firm (210 staff)
ScenarioWe make spend-management software for mid-sized companies and are exploring a reconciliation product. Our PM ran these 40-minute calls in August. Below are the relevant excerpts from each, lightly edited.
Call 1: Financial controller, logistics company (180 staff)“Close takes us eight working days and at least three of those are reconciliation. The bank feed drops transactions, so we tick and bash against statements in Excel.” Asked what she'd pay to fix it: “If it gave me two days back, it pays for itself. But I've heard that before.”
Call 2: Head of finance, SaaS company (95 staff)“Honestly close is fine. We moved to a proper ERP two years ago, it's five days and nobody's complaining.” Later: “The only annoying bit is chasing people for receipts, not reconciling.”
Call 3: Finance manager, retail chain (400 staff)“Card transactions are the nightmare. Hundreds of store cards, the statements come in as PDFs, someone re-keys them. We found £14k of duplicates last quarter.” She reconciles card spend in a shared spreadsheet with four people editing it.
Call 4: CFO, manufacturing company (250 staff)“We switched reconciliation tools last year and I regret it. The migration took four months, we lost all our matching rules, and half the team quietly went back to spreadsheets.” Asked if he'd switch again: “Not unless someone else does the migration for us. Never again otherwise.”
Call 5: Accountant, agency group (120 staff)“Intercompany is where it breaks. Three entities, money moving between them every week, and the balances never agree at month end. I spend a day and a half on it.”
Call 6: Finance director, charity (70 staff)“Close is fine for us. We're small, two of us, it takes four days and we know every transaction.” When asked about tools: “We'd not pay for anything new this year; budgets are frozen.”
Call 7: Controller, hospitality group (300 staff)“Payment processor payouts are the worst: they net off fees and refunds, so nothing matches the bank line. We reconcile by hand and it's always the last thing holding up close.”
Call 8: Financial controller, construction firm (210 staff)“Reconciliation is about 40% of my close. It's not hard, it's just volume, matching the same things every month.” On automation: “I'd want to see it match our messy data before I believed it. Every demo uses clean data.”
What a strong answer doesThe answer key the graders mark against

Themes with honest counts (five of eight describe reconciliation pain, from different sources), the two who say close is fine and the CFO's switching regret kept visible, and hypotheses about willingness to pay and switching cost labelled as such.

Critical failures (cap the score at 40)
  • Invents a quote
Case

Mid-level · v1.9 · anonymised real · B2B, finance

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit8 of 9 checks passed

The output is evidence-backed and actionable as-is. A PM would only need to add an explicit build trigger, which is a light edit.

Output

Month-end close discovery: findings summary

Bottom line

Reconciliation is a real, recurring pain for mid-sized finance teams, but these eight calls don't yet show which product to build or whether anyone will pay for it. I'd treat this as grounds for a further, more targeted round of research, not a build decision.

What we learned

1. Five of eight leads have a real reconciliation problem (calls 1, 3, 5, 7, 8). Two said close is fine (calls 2 and 6), and one has a tool and regrets it (call 4).

2. They describe different problems, not one. - Bank feeds dropping transactions, matched by hand in Excel (call 1) - Card statements arriving as PDFs and re-keyed, with £14k of duplicates in a quarter (call 3) - Intercompany balances that never agree, costing about 1.5 days (call 5) - Payment processor payouts that net off fees and refunds, and are always the last thing holding up close (call 7) - Sheer matching volume, about 40% of one controller's close (call 8)

No problem appears in more than one call, so we can't yet say which is most common or most valuable.

3. The time cost is meaningful but self-reported. Estimates ranged from 1.5 days to 3 of 8 days. These are recollections, not measurements.

4. Those who are fine tend to be smaller or better equipped. The two "close is fine" leads are the smallest companies in the sample (70 and 95 staff), and one already has a modern ERP. This is suggestive at best with eight calls. The pain sits mostly in the 180–400 staff range.

5. Adoption barriers are as prominent as the pain. - Call 4's failed migration (four months, lost matching rules, team back on spreadsheets) is a warning about switching costs. He would only switch again if someone else did the migration. - Call 8 won't believe automation works until he sees it handle messy data: "Every demo uses clean data." - Call 1 is sceptical: "I've heard that before."

6. Willingness to pay is essentially untested. Only call 1 gave a conditional answer ("if it gave me two days back, it pays for itself"), and she hedged it. Call 6 has a budget freeze. No one expressed urgency or asked about pricing.

7. An adjacent finding: call 2's only complaint was chasing receipts, which sits closer to our existing spend-management product than reconciliation does.

How confident can we be?

ClaimConfidence
Reconciliation is a meaningful time sink for some mid-sized teamsModerate
Pain varies by reconciliation typeModerate to high (consistent across calls)
Size of the time savingLow (self-reported, varied definitions)
Willingness to payLow (one hedged signal)
Which segment or use case to targetLow
Smaller or ERP-equipped firms are a weaker fitLow (three data points)

Other limitations: - Eight calls is a small sample, and we don't know how participants were recruited, so existing customers or those already interested may be overrepresented. - The excerpts are edited and one interviewer ran all the calls. - Hypothetical questions ("would you pay?") overstate real intent.

Suggested next steps

  1. Pick a wedge. Card reconciliation (call 3) is closest to our current product. Follow up with calls 3, 5 and 7 to size how often each problem occurs and what it costs.
  2. Test with real data. Ask call 8 and similar leads for anonymised, messy samples to see whether matching is feasible.
  3. Design for migration. Call 4 suggests onboarding and preserving matching rules may decide adoption. Treat this as a product requirement, not an afterthought.
  4. Test price and urgency. Use concrete pricing or a paid pilot, not hypotheticals.
  5. Widen the sample. Include firms of 150–500 staff and a mix of ERP and non-ERP setups, and record how each participant was recruited.

Grades and run details

Decision model 100 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly59%
  • passAddresses the actual decision93%
  • passRespects explicit constraints43%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims31%
  • passProduces the required deliverable78%
  • passKeeps dissent visible98%
  • passWeights behaviour over opinion98%
  • passSays how many sources support each finding65%
Run
Run
#1
API response time
18 s
Submitted
29 Sept 2026

Needs more than a quick edit3 of 9 checks passed

The output contains a false universal claim and lacks source counts per theme. A PM would need to re-check the call evidence before using it.

Output

Product Findings: Month-End Close & Reconciliation Opportunities

We spoke with eight finance leads at mid-sized companies (70–400 staff) to evaluate appetite for a reconciliation product. Here is what we learned, followed by our confidence assessment.

---

What We Learned

#### 1. Pain is Acute, but Fragmented by Niche Reconciliation is universally recognized as a bottleneck1, but the root cause varies heavily by business model: * Retail, Hospitality & Construction (High Volume/Messy Data): Companies dealing with multiple store cards (Call 3), complex payment processor payouts (Call 7), or high-volume line items (Call 8) spend up to 40% of their close cycle manually matching data in shared spreadsheets or re-keying PDF statements.3 * Agencies & Multi-Entity (Intercompany): Multi-entity structures suffer from recurring intercompany balancing discrepancies that defy easy automation4 (Call 5). * Logistics (Basic Bank Feeds): Basic bank feed drops force manual tick-and-bash routines in Excel (Call 1).

#### 2. Adoption Barriers and Skepticism Run High Even where pain is severe, market entry is fraught with buyer resistance: * Migration PTSD: Switching costs are massive. One CFO noted a four-month migration failure that destroyed custom matching rules and drove staff back to spreadsheets, refusing to switch again unless migration is fully managed (Call 4). * Deep Trust Deficit: Finance leads have been burned by over-promised automation and pristine demo environments. They demand proof on their own messy data before believing claims (Calls 1 & 8). * Budget Freezes & Apathy: Smaller teams (70–95 staff) with clean ERP setups or frozen budgets report no pressing need for new point solutions (Calls 2 & 6).

---

Confidence Assessment: Low-Moderate (Conditional)

  • Problem Validation (High Confidence): We are highly confident that reconciliation is tedious, manual, and a major driver of close delays in high-volume, multi-card, or multi-entity businesses.
  • Commercial Viability (Low-Moderate Confidence): Despite the pain, we cannot rely on a self-serve, software-only motion.2 High switching friction, deep skepticism toward AI/automation demos, and strict migration/implementation requirements mean buyers will churn or bounce without hands-on onboarding.

#### Recommendation Do not build a generic, self-serve reconciliation tool. If we proceed, the product must include a heavy-touch onboarding/migration service to overcome migration PTSD, alongside a "sandbox" feature allowing prospects to test matching rules on their actual messy data before purchasing.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimSubstantial rework

    Source: Call 2: Head of finance, SaaS company (95 staff)Honestly close is fine.

    What we’d changeSay most, not all: two of eight say close is fine. Keep that dissent visible rather than folding it into 'apathy'.

  2. 2Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changeScale the conclusion back to what eight calls show: adoption concerns from a few leads, not proof a self-serve product can't work. Give a confidence level per claim.

  3. 3Numbers wrongRedo the arithmeticTargeted repair

    Source: Call 8: Financial controller, construction firm (210 staff)Reconciliation is about 40% of my close.

    What we’d changeOnly Call 8 gives the 40% figure. Don't spread it across retail and hospitality.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    Source: Call 5: Accountant, agency group (120 staff)Intercompany is where it breaks.

    What we’d changeNothing in Call 5 says it defies automation. Report the day and a half it takes instead.

Claims the judge couldn’t find in the brief

  • Reconciliation is universally recognized as a bottleneck.
  • Finance leads have been burned by over-promised automation and pristine demo environments, and demand proof on their own messy data (Calls 1 & 8).

Grades and run details

Decision model 61 · LLM judge 3 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly27%
  • passAddresses the actual decision33%
  • passRespects explicit constraints28%
  • partialIdentifies material uncertainty53%
  • failAvoids unsupported claims67%
  • passProduces the required deliverable63%
  • passKeeps dissent visible88%
  • partialWeights behaviour over opinion21%
  • partialSays how many sources support each finding47%
Artefacts
Run
Run
#1
Time to output
7 s
Submitted
24 Sept 2026

Usable with a quick edit9 of 9 checks passed

The output is accurate, well-structured, and actionable with only minor wording changes needed.

Output

Findings: month-end reconciliation

These calls justify testing a focused reconciliation proposition, but do not yet justify building a broad product. Six participants described reconciliation pain or dissatisfaction with their tools. However, their problems differ substantially, and evidence of willingness to adopt or pay is weak.

What we learned

Reconciliation can materially delay close, but it is not a universal problem. The logistics controller reported at least three reconciliation days within an eight-day close. The agency accountant spends 1.5 days on intercompany balances. The construction controller estimated reconciliation at 40% of close, while the hospitality controller said processor payouts are the final blocker.

Conversely, the SaaS and charity participants explicitly said close works well. SaaS’s remaining frustration was receipt chasing, which should not be counted as reconciliation demand. These counterexamples argue against targeting all finance teams simply because they run a monthly close.

“Reconciliation” covers several distinct jobs.

ProblemEvidence from calls
Incomplete transaction dataLogistics: missing bank-feed transactions require checking statements in Excel
Card-statement processingRetail: PDF re-keying, shared spreadsheet coordination and £14k of duplicates found last quarter
Intercompany balancingAgency: transfers across three entities create mismatched balances
Processor payout reconciliationHospitality: fees and refunds obscure the relationship between payments and bank deposits
Repetitive matching at volumeConstruction: recurring manual work, described as time-consuming rather than difficult

These may require different data integrations, matching logic and workflows. Combining them into one demand signal would overstate support for any particular product.

Spreadsheets persist for different reasons. Logistics uses them to compensate for feed gaps; retail uses them to coordinate work; manufacturing returned to them after a failed tool migration. Spreadsheet use alone does not establish a desire to replace spreadsheets.

Trust and migration could determine adoption. The construction controller wants proof using messy, real data. The logistics controller is sceptical of promised savings. The manufacturing CFO made migration assistance a condition of switching after losing matching rules during a four-month migration. These are specific adoption requirements, although we cannot assume every buyer shares them.

How confident can we be?

Moderate confidence that meaningful problems exist in some of these organisations. Participants supplied concrete workflows, time estimates and one financial example. These remain self-reports: we have not observed the work or verified the figures. The £14k in duplicates does not establish losses incurred or savings our product could deliver.

Low confidence about prevalence, the best segment or a common solution. Eight exploratory calls provide useful hypotheses, not market estimates. Recruitment criteria and customer relationships are unspecified. Each industry appears only once, so differences cannot reliably be attributed to sector, company size or ERP maturity.

Very low confidence about commercial demand. One participant gave a conditional value statement after being asked about price, without naming a budget or committing to buy. Another explicitly ruled out new spending this year. No purchase, pilot commitment or switching decision is evidenced.

Implication for the product decision

Prioritise validation of card-spend reconciliation, given its proximity to our existing spend-management product, while recognising that only one call directly supports this use case.

Recruit more teams with that workflow, inspect actual statements and reconciliation records, and test a narrow prototype on messy data. Measure time saved, incorrect matches and exception-handling effort. Seek concrete pilot commitments and establish budget ownership and migration requirements before committing to build.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly46%
  • passAddresses the actual decision95%
  • passRespects explicit constraints41%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable71%
  • passKeeps dissent visible86%
  • passWeights behaviour over opinion87%
  • passSays how many sources support each finding53%
Run
Run
#1
Time to output
35 s
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyRightWrongRight
Sonnet 5.5 · API

All factual statements trace to the supplied calls, with no invented current-situation facts.

Gemini 3.5 Flash-Lite · Gemini

It asserts reconciliation is 'universally recognized as a bottleneck' even though Calls 2 and 6 describe close as fine, and it overstates Call 1/8 as 'burned by over-promised automation'.

GPT-6 Astra · ChatGPT

All facts and figures used are taken correctly from the supplied call excerpts with no invented current-state details.

Addresses the actual decisionMixedMixedRight
Sonnet 5.5 · API

It recommends more research but does not state what specific result or condition would trigger a build decision.

Gemini 3.5 Flash-Lite · Gemini

The call is only made at the end and does not state what result or condition would change it.

GPT-6 Astra · ChatGPT

The output clearly commits to testing a focused reconciliation proposition rather than building broadly, and states what validation results would change that call.

Identifies material uncertaintyRightWrongRight
Sonnet 5.5 · API

It names unknown WTP, segment, and time savings, and says how to resolve them in next steps.

Gemini 3.5 Flash-Lite · Gemini

It labels confidence low-moderate but does not name specific unknowns or how they would be resolved.

GPT-6 Astra · ChatGPT

It names prevalence, segment, willingness to pay, migration needs, and unverified self-reports as open questions and says how pilots and messy-data testing would resolve them.

Avoids unsupported claimsRightWrongRight
Sonnet 5.5 · API

Hypotheses and limitations are labelled as suggestive or untested, not as established fact.

Gemini 3.5 Flash-Lite · Gemini

It presents 'universally' and 'burned by over-promised automation' as established fact rather than as hypotheses.

GPT-6 Astra · ChatGPT

Interpretations and hypotheses are hedged or labelled as such, and confident claims are limited to those supported by the excerpts.

Weights behaviour over opinionRightWrongRight
Sonnet 5.5 · API

It flags self-reported recollections and hypothetical pricing as weaker than observed behaviours and workarounds.

Gemini 3.5 Flash-Lite · Gemini

It does not label themes by evidence type or flag stated preferences, such as Call 4's refusal to switch, as weaker than observed workarounds.

GPT-6 Astra · ChatGPT

It separates observed workarounds and past behaviour from stated intent and flags commercial demand as very low confidence.

Says how many sources support each findingRightWrongRight
Sonnet 5.5 · API

Findings consistently include source counts ('five of eight', 'two said', 'call 4').

Gemini 3.5 Flash-Lite · Gemini

Findings reference call numbers but do not say how many of the eight support each theme, and it uses an unsupported universal.

GPT-6 Astra · ChatGPT

Findings are anchored to named calls and explicit counts such as 'six participants', 'eight exploratory calls', 'only one call', and 'one participant/another', avoiding unsupported generalisation.

All got right 3

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

The output is a findings summary for the product team and is under the 600-word limit.

Gemini 3.5 Flash-Lite · Gemini

It delivers a findings summary under 600 words for the product team, with no apparent violation of stated constraints.

GPT-6 Astra · ChatGPT

It is a findings summary for the product team and is under the 600-word limit.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

The summary is complete, readable, and directly usable by the product team.

Gemini 3.5 Flash-Lite · Gemini

The required summary, confidence assessment and recommendation are present and readable, though some claims need correction.

GPT-6 Astra · ChatGPT

The summary is complete, usable, and would let a PM act with only light edits.

Keeps dissent visibleRightRightRight
Sonnet 5.5 · API

The two close-is-fine calls and the CFO's switching regret are kept visible throughout.

Gemini 3.5 Flash-Lite · Gemini

It keeps Call 2, Call 6, and Call 4's switching regret visible.

GPT-6 Astra · ChatGPT

The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 83% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.495.02None
2GPT-6.1 SolwithAPI94.490.02None
3GPT-6 LunawithAPI91.780.02None
4Sonnet 5.5withAPI94.465.02None
5Opus 5.5withClaude69.470.02None
6Gemini 3.8 FlashwithAPI52.855.02None
7Gemini 3.5 Flash-LitewithGemini58.330.02None

About the task

The PM job

Turning a stack of call transcripts into what we actually learned.

Why it matters

Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.

What good looks like

  • Quotes evidence for each theme and counts sources honestly
  • Weights what customers did above what they say they'd do
  • Keeps important dissent visible
  • Labels hypotheses as hypotheses
  • Says what the research cannot tell us

Deliberately not measured

  • Transcript clean-up
  • Persona illustration
Capability tested

Faithful synthesis of qualitative research

The failure we’re looking for

Invents customer consensus or loses important dissent

Grading

Decision model and LLM judge, calibrated against a blind PM review