Usable with a quick edit7 of 9 checks passed
The only error is a stray year in the title; the rest of the reasoning, evidence, and structure is solid and usable immediately after deleting '2026'.
Output
Reconciliation discovery: findings from 8 finance-lead calls (August 2026)3
Bottom line: Most of the finance leads we spoke to have real reconciliation pain. But it shows up as several different problems, not one, and the biggest barrier to adoption is distrust and switching cost. We have almost no evidence on willingness to pay. This justifies more targeted discovery. It doesn't yet justify a decision to build.
What we learned
- The pain is common but not universal. Six of eight say reconciliation is a significant part of close (Calls 1, 3, 4, 5, 7, 8).1 The two who said close is fine (2, 6) run the smallest companies, at 95 and 70 staff, and one of them has a modern ERP. That hints at a size threshold around 100+ staff, but it rests on two data points.
2. Where people gave numbers, the cost is large. - About 3 of 8 close days (1) - 1.5 days a month (5) - About 40% of close (8) - £14k of duplicates found in one quarter (3)
All of these are self-reported.
3. "Reconciliation" is at least five separate problems. Each call with pain named a different source: - Bank feed dropping transactions (1) - Card statements arriving as PDFs and re-keyed into a shared spreadsheet (3) - Intercompany balances that don't agree across entities (5) - Payment-processor payouts that net off fees and refunds, so nothing matches the bank line (7) - Sheer volume of repetitive matching (8)
No two calls named the same source. One "reconciliation product" would have to handle different matching logic for different segments. Card spend (3) sits closest to what we already do and could be a natural starting point, but it came up in only one call.
4. Switching cost and skepticism are the main barrier.2 - Call 4 switched reconciliation tools last year. Migration took four months, they lost their matching rules, and half the team drifted back to spreadsheets. He won't switch again unless someone else does the migration. - Call 8 won't believe a tool works until it handles their messy data, because demos always use clean data. - Call 1 has "heard that before."
That's three of eight, including people with heavy pain. Done-for-you migration and proof on the customer's own data look like requirements, not differentiators.
- One adjacent signal. Call 2's only complaint was chasing people for receipts. That's our existing spend-management territory, not reconciliation.
How confident can we be?
Moderate confidence: - Reconciliation takes a meaningful share of close time for mid-sized finance teams. - The underlying causes vary widely between companies. - Buyers are wary of switching tools and of vendor demos.
Low or no confidence: - Willingness to pay. Only one person answered the question, and only conditionally. One has a budget freeze. - Which problem to build for. Each pain source appears exactly once. - Size threshold. It rests on two calls. - Prevalence. We don't know how these eight were recruited, whether they're customers, or how representative they are.
Method limitations: - Eight calls, one interviewer, one month. - We're working from excerpts chosen as "relevant," so the full transcripts may shift the emphasis. - The pay question was open-ended and hypothetical.
What would raise confidence before a build decision
- Recruit 8–10 more finance leads at companies with 150–400 staff. Screen them by pain source so we learn which problem repeats.
- Test the card-spend starting point with existing customers. Ask how they reconcile card spend today. We may already hold the data needed to solve it.
- Test willingness to pay concretely. Use price anchors or offer a paid pilot rather than asking an open question.
- Run a matching test on one prospect's real, messy data. Call 8 set that bar, and it's the test that would convince the skeptics.
What a PM had to fix
1Numbers wrongRedo the arithmeticTargeted repair
Source: Call 4: CFO, manufacturing company (250 staff)
We switched reconciliation tools last year and I regret it.
What we’d changeRecount without Call 4, which is about a failed tool switch, not close time. It's five of eight.
2Hypothesis stated as factReframe it as a hypothesisQuick edit
What we’d changePresent it as a concern three people raised. Eight calls can't rank the barriers.
3Constraint missedRestore the constraintQuick edit
Source: Brief
Keep it under 600 words.
What we’d changeTrim it: it runs past the 600-word limit. And drop the year, which the brief doesn't give.
Check by check
Got wrong · 2
- Uses the supplied evidence correctlyThe output states the calls happened in August 2026, but the supplied context only says August without a year; this date is an invented fact not supported by the brief.
- Avoids unsupported claimsThe date 'August 2026' is presented as a confident fact without any basis in the supplied context.
Got right · 7
- Addresses the actual decisionThe output clearly recommends that the evidence doesn't yet justify building and that more targeted discovery is needed, and states what would change that call.
- Respects explicit constraintsThe output is a findings summary under 600 words for the product team, as requested.
- Identifies material uncertaintyIt names unknown willingness to pay, problem to build for, size threshold, and recruitment bias, and how to resolve them.
- Produces the required deliverableThe deliverable is a complete findings summary that a product team could act on with minimal edits.
- Keeps dissent visibleCalls 2 and 6 who said close is fine are explicitly mentioned and kept in view.
- Weights behaviour over opinionThe output flags self-reported numbers and hypothetical pay answers, and weights the actual migration experience (an observed behaviour) heavily.
- Says how many sources support each findingEvery finding is sized by number of calls (e.g., 'six of eight', 'three of eight', 'only one call').
Claims the judge couldn’t find in the brief
- The discovery calls took place in August 2026.
Grades and run details
Decision model 83 · LLM judge 8 of 10 checks
Decision model checks
- failUses the supplied evidence correctly22%
- passAddresses the actual decision95%
- passRespects explicit constraints22%
- passIdentifies material uncertainty100%
- partialAvoids unsupported claims18%
- passProduces the required deliverable45%
- passKeeps dissent visible49%
- passWeights behaviour over opinion87%
- passSays how many sources support each finding27%
Run
- Run
- #1
- Time to output
- 35 s
- Submitted
- 24 Sept 2026