Gemini 3.8 Flash
Google · released Sept 2026 · gemini-3.8-flash · effort medium (default)
What you can hand it, and what you’ll still have to catch
Dropped 1 Oct 2026, against whatever leads today.
Use it to find the angle, not to write the whole memo.
Gemini 3.8 Flash is good at spotting what a brief is really about. It found the mechanism hidden in the data on both 1000x cases, built roadmaps around their dependencies and a hard deadline, and wrote a customer call guide you could use with a quick edit. The trouble is everything around those insights: it fills gaps in the brief with facts of its own, and states its guesses as findings. That's why the LLM judge found only 1 of its 24 graded written outputs usable with a quick edit, and five of its 26 outputs made a critical mistake.
What you can hand it
- Customer research call guides. It asks what people actually did, gives each exec's theory a fair chance to fail, fixes a biased recruiting plan, and coaches first-time interviewers. Its freelancer guide was the one output usable with a quick edit.
- Finding the real mechanism. On both 1000x cases it found what the data was hiding (household coordination, workers carrying the app between jobs) instead of reaching for bigger adjectives.
- Sequencing a roadmap. It ordered work around every dependency, protected a hard processor deadline, and didn't plan on squads that haven't been hired yet.
- A first pass at a one-shot prototype. Its clinic rebooking prototype worked end to end; the receipt-capture one didn't, so try it before you share it.
What you’ll still have to catch
- Facts that aren't in the brief. Only about a quarter of its outputs used the supplied evidence correctly. It invented current systems in a PRD, a tooltip launch in an exec update, and "zero-CAC acquisition" in a strategy memo.
- Guesses stated as findings, in almost every output. It called a pricing test's mechanism proven, and treated founders' hypothesis as confirmed by interviews that didn't support it.
- Constraints it doesn't enforce. It gave a GO to a region with an open dietary incident, proposed sharing people's locations without their opt-in, and emailed meeting action items without the review the brief required.
- The working behind the plan. Roadmaps with no slack in the capacity sums, an experiment sized wrong with no plan for an inconclusive result, and discovery findings with no count of how many calls support them.
- Best setupGemini 3.8 Flash · API51.7#6 overall
- Against the leader−37.0GPT-6.1 Sol · API
Task by task
Its Combined score on each core task, across 26 outputs.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmap | Develop product strategy | Write a PRD | Customer research call guide | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specification | Analyse experiment results | 1000x an idea | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6.1 SolOpenAI | 89#1 | 87 | 87 | 69 | 94Best | 93 | 89 | 100Best | 88Best | 93 | 88 | 98Best | 80 | 88Best |
| Gemini 3.8 FlashGoogle | 52#6 | 58 | 52 | 38 | 73 | 51 | 76 | 56 | 47 | 37 | 48 | 48 | 41 | 47 |
How it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap58.2
- Customer research call guide72.7
- Experiment specification47.1
- Develop product strategy51.6
- One-shot prototype75.6
- Write a stakeholder update40.7
- Make the launch call47.3
- 1000x an idea48.1
- Write a PRD37.6
- Extract discovery insights51.3
- Challenge an idea48.1
- Analyse experiment results37.2
- Activation & onboarding review56.3
Gemini 3.8 Flash · API Other ranked setups. Hover a dot for its score and runs.
| Task | GPT-6.1 Sol · API | GPT-6 Astra · ChatGPT | GPT-6 Luna · API | Sonnet 5.5 · API | Opus 5.5 · Claude | Gemini 3.8 Flash · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|---|
| Build a roadmap | 86.5 (2 runs) | 87.8 (2 runs) | 77.2 (2 runs) | 95.2 (2 runs) | 71.6 (2 runs) | 58.2 (2 runs) | 17.2 (2 runs) |
| Customer research call guide | 93.5 (2 runs) | 85.5 (2 runs) | 61.1 (2 runs) | 85.9 (2 runs) | 74.2 (2 runs) | 72.7 (2 runs) | 38.0 (2 runs) |
| Experiment specification | 88.5 (2 runs) | 73.1 (2 runs) | 77.9 (2 runs) | 83.7 (2 runs) | 79.8 (2 runs) | 47.1 (2 runs) | 25.0 (2 runs) |
| Develop product strategy | 87.2 (2 runs) | 80.6 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 73.9 (2 runs) | 51.6 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 88.7 (2 runs) | 69.6 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 76.8 (2 runs) | 75.6 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 80.3 (2 runs) | 88.1 (2 runs) | 78.9 (2 runs) | 62.6 (2 runs) | 64.4 (2 runs) | 40.7 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 88.3 (2 runs) | 87.7 (2 runs) | 78.7 (2 runs) | 87.4 (2 runs) | 84.2 (2 runs) | 47.3 (2 runs) | 44.3 (2 runs) |
| 1000x an idea | 87.8 (2 runs) | 95.5 (2 runs) | 81.9 (2 runs) | 74.1 (2 runs) | 81.0 (2 runs) | 48.1 (2 runs) | 36.6 (2 runs) |
| Write a PRD | 69.0 (2 runs) | 95.2 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 64.0 (2 runs) | 37.6 (2 runs) | 42.8 (2 runs) |
| Extract discovery insights | 92.5 (2 runs) | 95.0 (2 runs) | 83.8 (2 runs) | 77.5 (2 runs) | 68.8 (2 runs) | 51.3 (2 runs) | 41.3 (2 runs) |
| Challenge an idea | 97.9 (2 runs) | 94.8 (2 runs) | 92.7 (2 runs) | 81.3 (2 runs) | 63.8 (2 runs) | 48.1 (2 runs) | 69.0 (2 runs) |
| Analyse experiment results | 92.8 (2 runs) | 97.5 (2 runs) | 91.6 (2 runs) | 79.7 (2 runs) | 61.9 (2 runs) | 37.2 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 100.0 (2 runs) | 97.5 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 71.3 (2 runs) | 56.3 (2 runs) | 58.8 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| 6 | Gemini 3.8 FlashwithAPIOpenRouter 2026-09 | 67.2 | 40.8 |
Strengths
Recurring failures
Bottom half on: Build a roadmap, Customer research call guide, Experiment specification, Develop product strategy, One-shot prototype, Write a stakeholder update, Make the launch call, 1000x an idea, Write a PRD, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review
By category
Gemini 3.8 Flash · API
Consistency
Typical API response time (app runs are timed by hand, so not comparable)
Published runs
26 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| 1000x an ideaFrom tip calculator to worker network | API | |
| 1000x an ideaSharing a shopping list, 1000x | API | |
| Activation & onboarding reviewFitness app first week | API | |
| Activation & onboarding reviewAnalytics tool losing users at setup | API | |
| Analyse experiment resultsThe underpowered onboarding test | API | |
| Analyse experiment resultsConversion up, retention down | API | |
| Build a roadmapA year of spend management, with a hard deadline | API | |
| Build a roadmapTwo squads, eight asks, one half | API | |
| Challenge an ideaThe CEO's embedded-payments bet | API | |
| Challenge an ideaAn AI SDR for small agencies | API | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | API | |
| Customer research call guideWhy are warehouses leaving? Two execs, two theories | API | |
| Develop product strategyHidden validation case | API | |
| Develop product strategySupply or demand for a stalled marketplace | API | |
| Experiment specificationBatching deliveries before peak season | API | |
| Experiment specificationShowing the delivery fee up front | API | |
| Extract discovery insightsHidden validation case | API | |
| Extract discovery insightsEight calls with finance teams | API | |
| Make the launch callGo/no-go for AI-drafted support replies | API | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | API | |
| One-shot prototypeClinic appointment rebooking | API | |
| One-shot prototypeExpense receipt capture | API | |
| Write a PRDAI triage for support tickets | API | |
| Write a PRDMeeting summaries with action items | API | |
| Write a stakeholder updateThe CEO asks why a metric dropped | API | |
| Write a stakeholder updateA quarterly business review deck, in brand | API |