Ratings / OpenAI

GPT-6 Luna

OpenAI · released 22 Sept 2026 · gpt-6-luna · effort medium (default)

GPT-6 LunawithAPIProvisional
23 runs · ±6.4 · Solid evidence

What you can hand it, and what you’ll still have to catch

Dropped 29 Sept 2026, against whatever leads today.

Sound judgement, but check its numbers.

Fourteen of Luna's 18 outputs were usable with at most a quick edit, and six were ready to use as written. Its calls held up: it applied agreed guardrails and launch gates as written, and challenged weak plans with evidence. Where it fell short was the working behind them: tables and ranges that don't reproduce, and economics left out of a strategy that asks for real money. Luna ran through the API rather than the ChatGPT app, so its results reflect the model on its own, without an app's tools or instructions.

What you can hand it

  • Launch calls and experiment readouts against agreed criteria. It held back a pricing variant that breached its guardrails, and didn't let keen agents or a booked training day override two failed launch gates.
  • Challenging a plan. Its AI SDR challenge was ready to use, and it found the payments model's central flaw: invoice value doesn't become processed volume at a 0.7% margin.
  • Replies to a CEO. It separated a confirmed usability problem from an unproven cause of a metric drop, and its QBR deck reported every OKR honestly, with an owner and a date on each decision.
  • One-shot prototypes. One was ready to share; the other needed its validation fixed.

What you’ll still have to catch

  • Tables and ranges that don't reproduce. Its rollout table mixed a later-period acceptance rate into a 1–28 October column and turned a pass into a fail, and its payments revenue range can't be rebuilt from its working.
  • Economics it skips. It committed the full £1.2m to a marketplace bet without showing how 6,000 bookings a month would pay for it.
  • Inconsistencies in the data it doesn't raise, and interview details it gets wrong, such as who said what.
  • PRD behaviour left undefined: what counts as an approved draft, and when an action's owner counts as confirmed.
  • Best setupGPT-6 Luna · API81.9provisional
  • Against the leader+8.7Opus 5.5 · Claude

Task by task

Its Combined score on each core task, across 23 outputs. Beside it, the mistakes a PM recorded in earlier reviews.

GPT-6 Luna's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6 LunaOpenAI82Partial–89Best6759838590769091918281
Opus 5.5Anthropic73#173746486717774776382646484
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

What you’ll need to check

The fixes you’ll make most often, each with a real example.

Verify or remove the claim

States a fact, figure, quote or current behaviour the source material doesn’t contain.

GPT-6 LunawithAPI on Supply or demand for a stalled marketplace
It wrote
Allocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.
What was missing

States a fact, figure, quote or current behaviour the source material doesn’t contain.

Verify or remove the claim · targeted repairTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.

Read the output

Tighten the test

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

GPT-6 LunawithAPI on AI triage for support tickets
It wrote
Approval sends the response through the existing support system; no response is sent before approval.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · targeted repairAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.

Read the output

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap–
  2. Customer research call guide59.3
  3. Experiment specification76.0
  4. Develop product strategy89.4leads
  5. One-shot prototype84.5
  6. Write a stakeholder update82.2
  7. Make the launch call80.5
  8. 1000x an idea90.9leads
  9. Write a PRD66.7
  10. Extract discovery insights82.5
  11. Challenge an idea90.6
  12. Analyse experiment results90.5
  13. Activation & onboarding review90.0

GPT-6 Luna · API Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
Prov.GPT-6 LunawithAPIOpenRouter 2026-0987.381.0

Strengths

    Leads: Develop product strategy, 1000x an idea

    Recurring failures

      Bottom half on: Customer research call guide, Experiment specification, Make the launch call, Write a PRD

      By category

      GPT-6 Luna · API

      Define78
      Discover71
      Design87
      Experiment83
      Challenge91
      Operate81

      Consistency

      31 s

      Typical API response time (app runs are timed by hand, so not comparable)

      Published runs

      23 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      1000x an ideaSharing a shopping list, 1000xAPI
      Activation & onboarding reviewAnalytics tool losing users at setupAPI
      Activation & onboarding reviewFitness app first weekAPI
      Analyse experiment resultsThe underpowered onboarding testAPI
      Analyse experiment resultsConversion up, retention downAPI
      Challenge an ideaAn AI SDR for small agenciesAPI
      Challenge an ideaThe CEO's embedded-payments betAPI
      Customer research call guideWould freelancers pay to stop chasing invoices?API
      Customer research call guideWhy are warehouses leaving? Two execs, two theoriesAPI
      Develop product strategySupply or demand for a stalled marketplaceAPI
      Develop product strategyHidden validation caseAPI
      Experiment specificationShowing the delivery fee up frontAPI
      Experiment specificationBatching deliveries before peak seasonAPI
      Extract discovery insightsEight calls with finance teamsAPI
      Extract discovery insightsHidden validation caseAPI
      Make the launch callGo/no-go for AI substitutions, from the rollout dataAPI
      Make the launch callGo/no-go for AI-drafted support repliesAPI
      One-shot prototypeExpense receipt captureAPI
      One-shot prototypeClinic appointment rebookingAPI
      Write a PRDAI triage for support ticketsAPI
      Write a PRDMeeting summaries with action itemsAPI
      Write a stakeholder updateA quarterly business review deck, in brandAPI
      Write a stakeholder updateThe CEO asks why a metric droppedAPI