Ratings / Google

Gemini 3.5 Flash-Lite

Google · released Jul 2026 · gemini-3.5-flash-lite · effort minimal (default)

Gemini 3.5 Flash-LitewithGeminiProvisional
19 runs · ±6.9 · Moderate evidence

What you can hand it, and what you’ll still have to catch

Dropped 25 Sept 2026, against whatever leads today.

Often the right direction, rarely the right numbers.

Flash-Lite usually lands on a defensible direction: hold the pricing variant, fix payroll first, don't launch on 2 November. The evidence it gives for that direction doesn't hold up. Two of its 18 outputs were usable, both of them prototypes. Ten needed starting again, mostly because of invented figures, misread tables, or safety constraints treated as met.

What you can hand it

  • One-shot prototypes. Both were rated Quick edit.
  • A first structure for a memo, as long as you redo the numbers yourself.

What you’ll still have to catch

  • Figures that aren't in the source: "~75% routing accuracy", "900,000+ interaction pairs", and an outbound share for an agency whose figure was never given.
  • Misread funnels and tables. It mixed funnel steps, spread one call's "40%" across three industries, and almost every figure in its rollout table was wrong.
  • Safety constraints it claims are met but aren't. It auto-emails action items despite an 8% speaker-label error rate, and calls "go with conditions" on two failed launch criteria.
  • Explanations presented as findings, almost everywhere.
  • Best setupGemini 3.5 Flash-Lite · Gemini48.1provisional
  • Against the leader−25.2Opus 5.5 · Claude

Task by task

Its Combined score on each core task, across 19 outputs. Beside it, the mistakes a PM recorded in earlier reviews.

Gemini 3.5 Flash-Lite's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
Opus 5.5Anthropic73#173746486717774776382646484
Gemini 3.5 Flash-LiteGoogle48Partial–5842–3949552353–675045
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

What you’ll need to check

The fixes you’ll make most often, each with a real example.

Reframe it as a hypothesis

Presents a plausible explanation or cause as if the evidence had established it.

Gemini 3.5 Flash-LitewithGemini on Supply or demand for a stalled marketplace
It wrote
The failure is entirely in initial activation and discovery.
The source (Research)

Students cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.

Reframe it as a hypothesis · start againPresent matching friction as the hypothesis the research suggests, not a proven diagnosis, and stage the spending so the board learns before it commits the full £1.2m.

Read the output

Verify or remove the claim

States a fact, figure, quote or current behaviour the source material doesn’t contain.

Gemini 3.5 Flash-LitewithGemini on AI triage for support tickets
It wrote
~75% (Manual/Rule-based)
The source (Volumes)

9,000 tickets a week; 38% are billing.

Verify or remove the claim · start againRemove the 75% routing baseline and the 900,000 interaction pairs: neither is in the brief. Measure the baseline before setting targets against it.

Read the output

Redo the arithmetic

Misreads the data, mixes up steps or denominators, or gets the sums wrong.

Gemini 3.5 Flash-LitewithGemini on Fitness app first week
It wrote
while 52% start the trial, only 19% complete three workouts in week one
The source (Funnel)

started trial 52% → first workout 48% → three workouts in week one 19%

Redo the arithmetic · start againRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.

Read the output

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap–
  2. Customer research call guide–
  3. Experiment specification23.1
  4. Develop product strategy57.7
  5. One-shot prototype49.4
  6. Write a stakeholder update50.2
  7. Make the launch call45.5
  8. 1000x an idea–
  9. Write a PRD41.7
  10. Extract discovery insights38.8
  11. Challenge an idea66.9
  12. Analyse experiment results52.5
  13. Activation & onboarding review55.0

Gemini 3.5 Flash-Lite · Gemini Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
Prov.Gemini 3.5 Flash-LitewithGeminiGemini 2026-0956.545.6

Strengths

    Recurring failures

      Bottom half on: Experiment specification, Develop product strategy, One-shot prototype, Write a stakeholder update, Make the launch call, Write a PRD, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review

      By category

      Gemini 3.5 Flash-Lite · Gemini

      Define50
      Discover39
      Design52
      Experiment38
      Challenge67
      Operate48

      Consistency

      16 s

      Typical time to output

      Published runs

      19 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      Activation & onboarding reviewFitness app first weekGemini
      Activation & onboarding reviewAnalytics tool losing users at setupGemini
      Analyse experiment resultsThe underpowered onboarding testGemini
      Analyse experiment resultsConversion up, retention downGemini
      Challenge an ideaAn AI SDR for small agenciesGemini
      Challenge an ideaThe CEO's embedded-payments betGemini
      Develop product strategyHidden validation caseGemini
      Develop product strategySupply or demand for a stalled marketplaceGemini
      Experiment specificationShowing the delivery fee up frontGemini
      Extract discovery insightsEight calls with finance teamsGemini
      Extract discovery insightsHidden validation caseGemini
      Make the launch callGo/no-go for AI-drafted support repliesGemini
      Make the launch callGo/no-go for AI substitutions, from the rollout dataGemini
      One-shot prototypeExpense receipt captureGemini
      One-shot prototypeClinic appointment rebookingGemini
      Write a PRDMeeting summaries with action itemsGemini
      Write a PRDAI triage for support ticketsGemini
      Write a stakeholder updateA quarterly business review deck, in brandGemini
      Write a stakeholder updateThe CEO asks why a metric droppedGemini