Ratings / Google

Gemini 3.8 Flash

Google · released Sept 2026 · gemini-3.8-flash · effort medium (default)

Gemini 3.8 FlashwithAPIRanked #6 overall
26 runs · ±5.9 · Solid evidence

What you can hand it, and what you’ll still have to catch

Dropped 1 Oct 2026, against whatever leads today.

Dropped

Use it to find the angle, not to write the whole memo.

Gemini 3.8 Flash is good at spotting what a brief is really about. It found the mechanism hidden in the data on both 1000x cases, built roadmaps around their dependencies and a hard deadline, and wrote a customer call guide you could use with a quick edit. The trouble is everything around those insights: it fills gaps in the brief with facts of its own, and states its guesses as findings. That's why the LLM judge found only 1 of its 24 graded written outputs usable with a quick edit, and five of its 26 outputs made a critical mistake.

What you can hand it

  • Customer research call guides. It asks what people actually did, gives each exec's theory a fair chance to fail, fixes a biased recruiting plan, and coaches first-time interviewers. Its freelancer guide was the one output usable with a quick edit.
  • Finding the real mechanism. On both 1000x cases it found what the data was hiding (household coordination, workers carrying the app between jobs) instead of reaching for bigger adjectives.
  • Sequencing a roadmap. It ordered work around every dependency, protected a hard processor deadline, and didn't plan on squads that haven't been hired yet.
  • A first pass at a one-shot prototype. Its clinic rebooking prototype worked end to end; the receipt-capture one didn't, so try it before you share it.

What you’ll still have to catch

  • Facts that aren't in the brief. Only about a quarter of its outputs used the supplied evidence correctly. It invented current systems in a PRD, a tooltip launch in an exec update, and "zero-CAC acquisition" in a strategy memo.
  • Guesses stated as findings, in almost every output. It called a pricing test's mechanism proven, and treated founders' hypothesis as confirmed by interviews that didn't support it.
  • Constraints it doesn't enforce. It gave a GO to a region with an open dietary incident, proposed sharing people's locations without their opt-in, and emailed meeting action items without the review the brief required.
  • The working behind the plan. Roadmaps with no slack in the capacity sums, an experiment sized wrong with no plan for an inconclusive result, and discovery findings with no count of how many calls support them.
  • Best setupGemini 3.8 Flash · API51.7#6 overall
  • Against the leader−37.0GPT-6.1 Sol · API

Task by task

Its Combined score on each core task, across 26 outputs.

Gemini 3.8 Flash's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapDevelop product strategyWrite a PRDCustomer research call guideExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationAnalyse experiment results1000x an ideaChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6.1 SolOpenAI89#187876994Best9389100Best88Best938898Best8088Best
Gemini 3.8 FlashGoogle52#658523873517656473748484147
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap58.2
  2. Customer research call guide72.7
  3. Experiment specification47.1
  4. Develop product strategy51.6
  5. One-shot prototype75.6
  6. Write a stakeholder update40.7
  7. Make the launch call47.3
  8. 1000x an idea48.1
  9. Write a PRD37.6
  10. Extract discovery insights51.3
  11. Challenge an idea48.1
  12. Analyse experiment results37.2
  13. Activation & onboarding review56.3

Gemini 3.8 Flash · API Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskGPT-6.1 Sol · APIGPT-6 Astra · ChatGPTGPT-6 Luna · APISonnet 5.5 · APIOpus 5.5 · ClaudeGemini 3.8 Flash · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap86.5 (2 runs)87.8 (2 runs)77.2 (2 runs)95.2 (2 runs)71.6 (2 runs)58.2 (2 runs)17.2 (2 runs)
Customer research call guide93.5 (2 runs)85.5 (2 runs)61.1 (2 runs)85.9 (2 runs)74.2 (2 runs)72.7 (2 runs)38.0 (2 runs)
Experiment specification88.5 (2 runs)73.1 (2 runs)77.9 (2 runs)83.7 (2 runs)79.8 (2 runs)47.1 (2 runs)25.0 (2 runs)
Develop product strategy87.2 (2 runs)80.6 (2 runs)89.4 (2 runs)83.5 (2 runs)73.9 (2 runs)51.6 (2 runs)57.7 (2 runs)
One-shot prototype88.7 (2 runs)69.6 (2 runs)84.5 (2 runs)92.9 (2 runs)76.8 (2 runs)75.6 (2 runs)49.4 (2 runs)
Write a stakeholder update80.3 (2 runs)88.1 (2 runs)78.9 (2 runs)62.6 (2 runs)64.4 (2 runs)40.7 (2 runs)50.2 (2 runs)
Make the launch call88.3 (2 runs)87.7 (2 runs)78.7 (2 runs)87.4 (2 runs)84.2 (2 runs)47.3 (2 runs)44.3 (2 runs)
1000x an idea87.8 (2 runs)95.5 (2 runs)81.9 (2 runs)74.1 (2 runs)81.0 (2 runs)48.1 (2 runs)36.6 (2 runs)
Write a PRD69.0 (2 runs)95.2 (2 runs)66.7 (2 runs)75.6 (2 runs)64.0 (2 runs)37.6 (2 runs)42.8 (2 runs)
Extract discovery insights92.5 (2 runs)95.0 (2 runs)83.8 (2 runs)77.5 (2 runs)68.8 (2 runs)51.3 (2 runs)41.3 (2 runs)
Challenge an idea97.9 (2 runs)94.8 (2 runs)92.7 (2 runs)81.3 (2 runs)63.8 (2 runs)48.1 (2 runs)69.0 (2 runs)
Analyse experiment results92.8 (2 runs)97.5 (2 runs)91.6 (2 runs)79.7 (2 runs)61.9 (2 runs)37.2 (2 runs)52.5 (2 runs)
Activation & onboarding review100.0 (2 runs)97.5 (2 runs)90.0 (2 runs)57.5 (2 runs)71.3 (2 runs)56.3 (2 runs)58.8 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
6Gemini 3.8 FlashwithAPIOpenRouter 2026-0967.240.8

Strengths

    Recurring failures

      Bottom half on: Build a roadmap, Customer research call guide, Experiment specification, Develop product strategy, One-shot prototype, Write a stakeholder update, Make the launch call, 1000x an idea, Write a PRD, Extract discovery insights, Challenge an idea, Analyse experiment results, Activation & onboarding review

      By category

      Gemini 3.8 Flash · API

      Define49
      Discover62
      Design66
      Experiment42
      Challenge48
      Operate44

      Consistency

      32 s

      Typical API response time (app runs are timed by hand, so not comparable)

      Published runs

      26 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      1000x an ideaFrom tip calculator to worker networkAPI
      1000x an ideaSharing a shopping list, 1000xAPI
      Activation & onboarding reviewFitness app first weekAPI
      Activation & onboarding reviewAnalytics tool losing users at setupAPI
      Analyse experiment resultsThe underpowered onboarding testAPI
      Analyse experiment resultsConversion up, retention downAPI
      Build a roadmapA year of spend management, with a hard deadlineAPI
      Build a roadmapTwo squads, eight asks, one halfAPI
      Challenge an ideaThe CEO's embedded-payments betAPI
      Challenge an ideaAn AI SDR for small agenciesAPI
      Customer research call guideWould freelancers pay to stop chasing invoices?API
      Customer research call guideWhy are warehouses leaving? Two execs, two theoriesAPI
      Develop product strategyHidden validation caseAPI
      Develop product strategySupply or demand for a stalled marketplaceAPI
      Experiment specificationBatching deliveries before peak seasonAPI
      Experiment specificationShowing the delivery fee up frontAPI
      Extract discovery insightsHidden validation caseAPI
      Extract discovery insightsEight calls with finance teamsAPI
      Make the launch callGo/no-go for AI-drafted support repliesAPI
      Make the launch callGo/no-go for AI substitutions, from the rollout dataAPI
      One-shot prototypeClinic appointment rebookingAPI
      One-shot prototypeExpense receipt captureAPI
      Write a PRDAI triage for support ticketsAPI
      Write a PRDMeeting summaries with action itemsAPI
      Write a stakeholder updateThe CEO asks why a metric droppedAPI
      Write a stakeholder updateA quarterly business review deck, in brandAPI