Tasks / Discover

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Respects explicit constraints100% pass
    It is a findings summary for the product team and is under the 600-word limit.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  2. Produces the required deliverable100% pass
    The summary is complete, usable, and would let a PM act with only light edits.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams
  3. Keeps dissent visible100% pass
    The SaaS and charity close-is-fine views and the manufacturing CFO's switching regret are kept visible.
    GPT-6 Astra · ChatGPT · Eight calls with finance teams

Where it slips

  1. Says how many sources support each finding52% pass
    Findings reference call numbers but do not say how many of the eight support each theme, and it uses an unsupported universal.
    Gemini 3.5 Flash-Lite · Gemini · Eight calls with finance teams
  2. Avoids unsupported claims56% pass
    The bottom line asserts 'Adoption will depend on handling messy data and reducing switching effort' as an established forecast, though the evidence only suggests it.
    GPT-6.1 Sol · API · Eight calls with finance teams
  3. Uses the supplied evidence correctly65% pass
    The output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
    Opus 5.5 · Claude · Eight calls with finance teams

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Synthesise the eight discovery calls below with finance leads about month-end close. Write a findings summary for the product team, who are deciding whether to build a reconciliation product: what we learned, and how confident we can be in it. Keep it under 600 words.

ScenarioWe make spend-management software for mid-sized companies and are exploring a reconciliation product. Our PM ran these 40-minute calls in August. Below are the relevant excerpts from each, lightly edited.
Call 1: Financial controller, logistics company (180 staff)“Close takes us eight working days and at least three of those are reconciliation. The bank feed drops transactions, so we tick and bash against statements in Excel.” Asked what she'd pay to fix it: “If it gave me two days back, it pays for itself. But I've heard that before.”
Call 2: Head of finance, SaaS company (95 staff)“Honestly close is fine. We moved to a proper ERP two years ago, it's five days and nobody's complaining.” Later: “The only annoying bit is chasing people for receipts, not reconciling.”
Call 3: Finance manager, retail chain (400 staff)“Card transactions are the nightmare. Hundreds of store cards, the statements come in as PDFs, someone re-keys them. We found £14k of duplicates last quarter.” She reconciles card spend in a shared spreadsheet with four people editing it.
Call 4: CFO, manufacturing company (250 staff)“We switched reconciliation tools last year and I regret it. The migration took four months, we lost all our matching rules, and half the team quietly went back to spreadsheets.” Asked if he'd switch again: “Not unless someone else does the migration for us. Never again otherwise.”
Call 5: Accountant, agency group (120 staff)“Intercompany is where it breaks. Three entities, money moving between them every week, and the balances never agree at month end. I spend a day and a half on it.”
Call 6: Finance director, charity (70 staff)“Close is fine for us. We're small, two of us, it takes four days and we know every transaction.” When asked about tools: “We'd not pay for anything new this year; budgets are frozen.”
Call 7: Controller, hospitality group (300 staff)“Payment processor payouts are the worst: they net off fees and refunds, so nothing matches the bank line. We reconcile by hand and it's always the last thing holding up close.”
Call 8: Financial controller, construction firm (210 staff)“Reconciliation is about 40% of my close. It's not hard, it's just volume, matching the same things every month.” On automation: “I'd want to see it match our messy data before I believed it. Every demo uses clean data.”
What a strong answer does

Themes with honest counts (five of eight describe reconciliation pain, from different sources), the two who say close is fine and the CFO's switching regret kept visible, and hypotheses about willingness to pay and switching cost labelled as such.

Critical failures (cap the score)
  • Invents a quote
Case

v1.7 · anonymised real · B2B, finance

Usable with a quick edit7 of 9 checks passed

The only error is a stray year in the title; the rest of the reasoning, evidence, and structure is solid and usable immediately after deleting '2026'.

Output

Reconciliation discovery: findings from 8 finance-lead calls (August 2026)3

Bottom line: Most of the finance leads we spoke to have real reconciliation pain. But it shows up as several different problems, not one, and the biggest barrier to adoption is distrust and switching cost. We have almost no evidence on willingness to pay. This justifies more targeted discovery. It doesn't yet justify a decision to build.

What we learned

  1. The pain is common but not universal. Six of eight say reconciliation is a significant part of close (Calls 1, 3, 4, 5, 7, 8).1 The two who said close is fine (2, 6) run the smallest companies, at 95 and 70 staff, and one of them has a modern ERP. That hints at a size threshold around 100+ staff, but it rests on two data points.

2. Where people gave numbers, the cost is large. - About 3 of 8 close days (1) - 1.5 days a month (5) - About 40% of close (8) - £14k of duplicates found in one quarter (3)

All of these are self-reported.

3. "Reconciliation" is at least five separate problems. Each call with pain named a different source: - Bank feed dropping transactions (1) - Card statements arriving as PDFs and re-keyed into a shared spreadsheet (3) - Intercompany balances that don't agree across entities (5) - Payment-processor payouts that net off fees and refunds, so nothing matches the bank line (7) - Sheer volume of repetitive matching (8)

No two calls named the same source. One "reconciliation product" would have to handle different matching logic for different segments. Card spend (3) sits closest to what we already do and could be a natural starting point, but it came up in only one call.

4. Switching cost and skepticism are the main barrier.2 - Call 4 switched reconciliation tools last year. Migration took four months, they lost their matching rules, and half the team drifted back to spreadsheets. He won't switch again unless someone else does the migration. - Call 8 won't believe a tool works until it handles their messy data, because demos always use clean data. - Call 1 has "heard that before."

That's three of eight, including people with heavy pain. Done-for-you migration and proof on the customer's own data look like requirements, not differentiators.

  1. One adjacent signal. Call 2's only complaint was chasing people for receipts. That's our existing spend-management territory, not reconciliation.

How confident can we be?

Moderate confidence: - Reconciliation takes a meaningful share of close time for mid-sized finance teams. - The underlying causes vary widely between companies. - Buyers are wary of switching tools and of vendor demos.

Low or no confidence: - Willingness to pay. Only one person answered the question, and only conditionally. One has a budget freeze. - Which problem to build for. Each pain source appears exactly once. - Size threshold. It rests on two calls. - Prevalence. We don't know how these eight were recruited, whether they're customers, or how representative they are.

Method limitations: - Eight calls, one interviewer, one month. - We're working from excerpts chosen as "relevant," so the full transcripts may shift the emphasis. - The pay question was open-ended and hypothetical.

What would raise confidence before a build decision

  • Recruit 8–10 more finance leads at companies with 150–400 staff. Screen them by pain source so we learn which problem repeats.
  • Test the card-spend starting point with existing customers. Ask how they reconcile card spend today. We may already hold the data needed to solve it.
  • Test willingness to pay concretely. Use price anchors or offer a paid pilot rather than asking an open question.
  • Run a matching test on one prospect's real, messy data. Call 8 set that bar, and it's the test that would convince the skeptics.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Call 4: CFO, manufacturing company (250 staff)We switched reconciliation tools last year and I regret it.

    What we’d changeRecount without Call 4, which is about a failed tool switch, not close time. It's five of eight.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent it as a concern three people raised. Eight calls can't rank the barriers.

  3. 3Constraint missedRestore the constraintQuick edit

    Source: BriefKeep it under 600 words.

    What we’d changeTrim it: it runs past the 600-word limit. And drop the year, which the brief doesn't give.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyThe output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
  • Avoids unsupported claimsThe date 'August 2026' is presented as a confident fact without any basis in the supplied context.
Got right · 7
  • Addresses the actual decisionThe output clearly recommends that the evidence doesn't yet justify building and that more targeted discovery is needed, and states what would change that call.
  • Respects explicit constraintsThe output is a findings summary under 600 words for the product team, as requested.
  • Identifies material uncertaintyIt names unknown willingness to pay, problem to build for, size threshold, and recruitment bias, and how to resolve them.
  • Produces the required deliverableThe deliverable is a complete findings summary that a product team could act on with minimal edits.
  • Keeps dissent visibleCalls 2 and 6 who said close is fine are explicitly mentioned and kept in view.
  • Weights behaviour over opinionThe output flags self-reported numbers and hypothetical pay answers, and weights the actual migration experience (an observed behaviour) heavily.
  • Says how many sources support each findingEvery finding is sized by number of calls (e.g., 'six of eight', 'three of eight', 'only one call').

Claims the judge couldn’t find in the brief

  • The discovery calls took place in August 2026.

Grades and run details

Decision model 83 · LLM judge 8 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly22%
  • passAddresses the actual decision95%
  • passRespects explicit constraints22%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims18%
  • passProduces the required deliverable45%
  • passKeeps dissent visible49%
  • passWeights behaviour over opinion87%
  • passSays how many sources support each finding27%
Run
Run
#1
Time to output
35 s
Submitted
24 Sept 2026

Usable with a quick edit9 of 9 checks passed

The output is evidence-accurate, well-structured, and actionable with no need to re-check facts or redo reasoning.

Output

Findings summary

The pain is real, but it is not universal—and it is concentrated in specific reconciliation workflows. Five of the eight interviewees described reconciliation as a meaningful close burden. Their problems were concrete: dropped bank-feed transactions and manual matching (Call 1); high-volume card statements re-keyed from PDFs (Call 3); intercompany balances (Call 5); processor payouts obscured by fees and refunds (Call 7); and repetitive matching at scale (Call 8). One reported finding £14k in duplicate transactions. These examples suggest the opportunity may be less “speed up month-end close” and more “handle particular messy, high-volume matching jobs reliably.”

Company context seems to matter.1 The clearest pain came from organisations with volume, multiple entities, many cards, or complex payment flows. Two interviewees said close was fine: one had a well-established ERP and mainly struggled with receipts; another was small, knew every transaction, and had frozen budgets. A reconciliation product is therefore unlikely to be equally valuable to every mid-sized company.

Trust and implementation are central adoption barriers. One finance leader regretted a tool switch after a four-month migration, lost matching rules, and a return to spreadsheets. Another would not trust automation until it worked on messy real data. A product that matches well in a demo may still fail if it cannot preserve existing rules, handle exceptions, and prove accuracy on customers’ own data.

Willingness to pay is unproven. One controller said saving two days could justify the cost, but explicitly expressed scepticism based on past promises. Another prospect would not buy this year. These calls reveal value hypotheses, not validated demand or pricing.

Implications for the product team

Explore a narrow, workflow-led wedge rather than assuming a broad reconciliation platform is needed. Card-spend reconciliation may fit our existing product particularly well; processor payouts and intercompany are other distinct opportunities, but may require different data and capabilities. Test which segment has both acute pain and access to usable data. Prioritise messy-data performance, explainable matches and exceptions, and low-friction migration or coexistence with spreadsheets.

Confidence

  • Moderate confidence that reconciliation is a significant pain for some complex, high-volume finance teams: several independent, specific examples support this.
  • Low-to-moderate confidence on which workflow or customer segment is the best product wedge: each use case appears in few calls.
  • Low confidence on market-wide prevalence, willingness to pay, achievable time savings, or product feasibility: eight qualitative calls cannot establish these, and no prototype or pricing was tested.

Treat this as a reason to run targeted workflow and data-validation research—not yet as evidence to commit to a broad build.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    What we’d changeCompany context matters.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll factual claims about the current situation trace directly to specific calls in the supplied context with no invented facts.
  • Addresses the actual decisionThe output commits to not committing to a broad build now and running targeted validation, and it says what to test (segment with acute pain and usable data).
  • Respects explicit constraintsThe output is under 600 words, addressed to the product team, and in the requested findings-summary form.
  • Identifies material uncertaintyIt names the key unknowns (willingness to pay, time savings, workflow/segment, prevalence, feasibility) and says targeted workflow and data-validation research would resolve them.
  • Avoids unsupported claimsHypotheses and interpretations are hedged with 'suggest', 'may', 'seems', 'unlikely', and 'unproven'; no unsupported confident causal claims are presented as fact.
  • Produces the required deliverableThe findings summary is complete, actionable, and usable for the product team without major gaps.
  • Keeps dissent visibleThe dissenting interviews (Calls 2 and 6, and the switch-regret in Call 4) are explicitly included.
  • Weights behaviour over opinionIt highlights actual behaviour (Excel tick-and-bash, PDF re-keying, shared spreadsheet, migration, manual processor reconciliation) and flags willingness to pay as unproven stated intent.
  • Says how many sources support each findingFindings are sized with 'five of eight', 'two interviewees', 'one finance leader', and 'another' rather than generalising to everyone.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly65%
  • passAddresses the actual decision46%
  • passRespects explicit constraints45%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims39%
  • passProduces the required deliverable85%
  • passKeeps dissent visible90%
  • passWeights behaviour over opinion94%
  • passSays how many sources support each finding77%
Run
Run
#1
API response time
11 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 85% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT97.295.02None
2GPT-6.1 SolwithAPI94.490.02None
3GPT-6 LunawithAPI88.980.02None
4Sonnet 5.5withAPI94.465.02None
5Opus 5.5withClaude75.070.02None
6Gemini 3.5 Flash-LitewithGemini52.830.02None

About the task

The PM job

Turning a stack of call transcripts into what we actually learned.

Why it matters

Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.

What good looks like

  • Quotes evidence for each theme and counts sources honestly
  • Weights what customers did above what they say they'd do
  • Keeps important dissent visible
  • Labels hypotheses as hypotheses
  • Says what the research cannot tell us

Deliberately not measured

  • Transcript clean-up
  • Persona illustration
Capability tested

Faithful synthesis of qualitative research

The failure we’re looking for

Invents customer consensus or loses important dissent

Grading

Decision model and LLM judge, calibrated against a blind PM review