Gemini 3.5 Flash-Lite
Google · released Jul 2026 · gemini-3.5-flash-lite · effort minimal (default)
What you can hand it, and what you’ll still have to catch
Dropped 25 Sept 2026, against whatever leads today.
Often the right direction, rarely the right numbers.
Flash-Lite usually lands on a defensible direction: hold the pricing variant, fix payroll first, don't launch on 2 November. The evidence it gives for that direction doesn't hold up. Two of its 18 outputs were usable, both of them prototypes. Ten needed starting again, mostly because of invented figures, misread tables, or safety constraints treated as met.
What you can hand it
- One-shot prototypes. Both were rated Quick edit.
- A first structure for a memo, as long as you redo the numbers yourself.
What you’ll still have to catch
- Figures that aren't in the source: "~75% routing accuracy", "900,000+ interaction pairs", and an outbound share for an agency whose figure was never given.
- Misread funnels and tables. It mixed funnel steps, spread one call's "40%" across three industries, and almost every figure in its rollout table was wrong.
- Safety constraints it claims are met but aren't. It auto-emails action items despite an 8% speaker-label error rate, and calls "go with conditions" on two failed launch criteria.
- Explanations presented as findings, almost everywhere.
- Best setupGemini 3.5 Flash-Lite · Gemini48.1provisional
- Against the leader−25.2Opus 5.5 · Claude
Task by task
Its Combined score on each core task, across 19 outputs. Beside it, the mistakes a PM recorded in earlier reviews.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
| Gemini 3.5 Flash-LiteGoogle | 48Partial | – | 58 | 42 | – | 39 | 49 | 55 | 23 | 53 | – | 67 | 50 | 45 |
- Hypothesis stated as fact13 of 19Extract discovery insights, Develop product strategy, Write a stakeholder update
- Invented evidence11 of 19Write a stakeholder update, Develop product strategy, Extract discovery insights
- Numbers wrong6 of 19Challenge an idea, Activation & onboarding review, Make the launch call
- Constraint missed4 of 19Write a PRD, Make the launch call, Develop product strategy
- Contradiction missed4 of 19Develop product strategy, Make the launch call, Challenge an idea
- Test or gate too weak2 of 19Write a PRD, Challenge an idea
- Artefact unusable2 of 19One-shot prototype
- Other1 of 19Analyse experiment results
What you’ll need to check
The fixes you’ll make most often, each with a real example.
Reframe it as a hypothesis
Presents a plausible explanation or cause as if the evidence had established it.
The failure is entirely in initial activation and discovery.
Students cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Reframe it as a hypothesis · start againPresent matching friction as the hypothesis the research suggests, not a proven diagnosis, and stage the spending so the board learns before it commits the full £1.2m.
Read the outputVerify or remove the claim
States a fact, figure, quote or current behaviour the source material doesn’t contain.
~75% (Manual/Rule-based)
9,000 tickets a week; 38% are billing.
Verify or remove the claim · start againRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.
Read the outputRedo the arithmetic
Misreads the data, mixes up steps or denominators, or gets the sums wrong.
while 52% start the trial, only 19% complete three workouts in week one
started trial 52% → first workout 48% → three workouts in week one 19%
Redo the arithmetic · start againRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.
Read the outputHow it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap–
- Customer research call guide–
- Experiment specification23.1
- Develop product strategy57.7
- One-shot prototype49.4
- Write a stakeholder update50.2
- Make the launch call45.5
- 1000x an idea–
- Write a PRD41.7
- Extract discovery insights38.8
- Challenge an idea66.9
- Analyse experiment results52.5
- Activation & onboarding review55.0
Gemini 3.5 Flash-Lite · Gemini Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| Prov. | Gemini 3.5 Flash-LitewithGeminiGemini 2026-09 | 56.5 | 45.6 |
Strengths
By category
Gemini 3.5 Flash-Lite · Gemini
Consistency
Typical time to output
Published runs
19 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| Activation & onboarding reviewFitness app first week | Gemini | |
| Activation & onboarding reviewAnalytics tool losing users at setup | Gemini | |
| Analyse experiment resultsThe underpowered onboarding test | Gemini | |
| Analyse experiment resultsConversion up, retention down | Gemini | |
| Challenge an ideaAn AI SDR for small agencies | Gemini | |
| Challenge an ideaThe CEO's embedded-payments bet | Gemini | |
| Develop product strategyHidden validation case | Gemini | |
| Develop product strategySupply or demand for a stalled marketplace | Gemini | |
| Experiment specificationShowing the delivery fee up front | Gemini | |
| Extract discovery insightsEight calls with finance teams | Gemini | |
| Extract discovery insightsHidden validation case | Gemini | |
| Make the launch callGo/no-go for AI-drafted support replies | Gemini | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | Gemini | |
| One-shot prototypeExpense receipt capture | Gemini | |
| One-shot prototypeClinic appointment rebooking | Gemini | |
| Write a PRDMeeting summaries with action items | Gemini | |
| Write a PRDAI triage for support tickets | Gemini | |
| Write a stakeholder updateA quarterly business review deck, in brand | Gemini | |
| Write a stakeholder updateThe CEO asks why a metric dropped | Gemini |