GPT-6 Luna
OpenAI · released 22 Sept 2026 · gpt-6-luna · effort medium (default)
What you can hand it, and what you’ll still have to catch
Dropped 29 Sept 2026, against whatever leads today.
Sound judgement, but check its numbers.
Fourteen of Luna's 18 outputs were usable with at most a quick edit, and six were ready to use as written. Its calls held up: it applied agreed guardrails and launch gates as written, and challenged weak plans with evidence. Where it fell short was the working behind them: tables and ranges that don't reproduce, and economics left out of a strategy that asks for real money. Luna ran through the API rather than the ChatGPT app, so its results reflect the model on its own, without an app's tools or instructions.
What you can hand it
- Launch calls and experiment readouts against agreed criteria. It held back a pricing variant that breached its guardrails, and didn't let keen agents or a booked training day override two failed launch gates.
- Challenging a plan. Its AI SDR challenge was ready to use, and it found the payments model's central flaw: invoice value doesn't become processed volume at a 0.7% margin.
- Replies to a CEO. It separated a confirmed usability problem from an unproven cause of a metric drop, and its QBR deck reported every OKR honestly, with an owner and a date on each decision.
- One-shot prototypes. One was ready to share; the other needed its validation fixed.
What you’ll still have to catch
- Tables and ranges that don't reproduce. Its rollout table mixed a later-period acceptance rate into a 1–28 October column and turned a pass into a fail, and its payments revenue range can't be rebuilt from its working.
- Economics it skips. It committed the full £1.2m to a marketplace bet without showing how 6,000 bookings a month would pay for it.
- Inconsistencies in the data it doesn't raise, and interview details it gets wrong, such as who said what.
- PRD behaviour left undefined: what counts as an approved draft, and when an action's owner counts as confirmed.
- Best setupGPT-6 Luna · API81.9provisional
- Against the leader+8.7Opus 5.5 · Claude
Task by task
Its Combined score on each core task, across 23 outputs. Beside it, the mistakes a PM recorded in earlier reviews.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6 LunaOpenAI | 82Partial | – | 89Best | 67 | 59 | 83 | 85 | 90 | 76 | 90 | 91 | 91 | 82 | 81 |
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
- Invented evidence3 of 23Develop product strategy, Make the launch call, Extract discovery insights
- Test or gate too weak2 of 23Write a PRD, Extract discovery insights
- Contradiction missed1 of 23Develop product strategy
- Other1 of 23Challenge an idea
What you’ll need to check
The fixes you’ll make most often, each with a real example.
Verify or remove the claim
States a fact, figure, quote or current behaviour the source material doesn’t contain.
Allocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.
States a fact, figure, quote or current behaviour the source material doesn’t contain.
Verify or remove the claim · targeted repairTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.
Read the outputTighten the test
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Approval sends the response through the existing support system; no response is sent before approval.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · targeted repairAn agent’s explicit “Approve and send” action authorises only the exact response displayed, including their edits. Regeneration or any subsequent change requires fresh approval. Enforce this in the send service and prevent duplicate sends on retries. If drafting fails, agents can compose and send manually. AI-generated drafts must use approved noncommittal refund wording; agents may add refund commitments under existing policy and explicitly approve the final response.
Read the outputHow it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap–
- Customer research call guide59.3
- Experiment specification76.0
- Develop product strategy89.4leads
- One-shot prototype84.5
- Write a stakeholder update82.2
- Make the launch call80.5
- 1000x an idea90.9leads
- Write a PRD66.7
- Extract discovery insights82.5
- Challenge an idea90.6
- Analyse experiment results90.5
- Activation & onboarding review90.0
GPT-6 Luna · API Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| Prov. | GPT-6 LunawithAPIOpenRouter 2026-09 | 87.3 | 81.0 |
Recurring failures
Bottom half on: Customer research call guide, Experiment specification, Make the launch call, Write a PRD
By category
GPT-6 Luna · API
Consistency
Typical API response time (app runs are timed by hand, so not comparable)
Published runs
23 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| 1000x an ideaSharing a shopping list, 1000x | API | |
| Activation & onboarding reviewAnalytics tool losing users at setup | API | |
| Activation & onboarding reviewFitness app first week | API | |
| Analyse experiment resultsThe underpowered onboarding test | API | |
| Analyse experiment resultsConversion up, retention down | API | |
| Challenge an ideaAn AI SDR for small agencies | API | |
| Challenge an ideaThe CEO's embedded-payments bet | API | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | API | |
| Customer research call guideWhy are warehouses leaving? Two execs, two theories | API | |
| Develop product strategySupply or demand for a stalled marketplace | API | |
| Develop product strategyHidden validation case | API | |
| Experiment specificationShowing the delivery fee up front | API | |
| Experiment specificationBatching deliveries before peak season | API | |
| Extract discovery insightsEight calls with finance teams | API | |
| Extract discovery insightsHidden validation case | API | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | API | |
| Make the launch callGo/no-go for AI-drafted support replies | API | |
| One-shot prototypeExpense receipt capture | API | |
| One-shot prototypeClinic appointment rebooking | API | |
| Write a PRDAI triage for support tickets | API | |
| Write a PRDMeeting summaries with action items | API | |
| Write a stakeholder updateA quarterly business review deck, in brand | API | |
| Write a stakeholder updateThe CEO asks why a metric dropped | API |