GPT-6 Astra
OpenAI · gpt-6-astra · effort medium (default)
What you can hand it, and what you’ll still have to catch
Dropped 25 Sept 2026, against whatever leads today.
Handles most PM work well, but sometimes won't commit.
Fifteen of Astra's 18 outputs were usable with at most a quick edit, and six were ready to use as written. It keeps what the evidence shows apart from what it suggests, better than any setup so far. When it falls short, it hedges where a decision was needed, or it proposes a test that couldn't settle the question.
What you can hand it
- Discovery synthesis. Both cases were rated Ready to use, with confidence matched to the evidence.
- Strategy memos and challenges to a plan. It separates evidence from inference and shows its working.
- Go/no-go calls against agreed criteria. It applies each gate as written, and doesn't let human review excuse a failure.
- PRDs where the safety constraint is the hard part.
What you’ll still have to catch
- Hedging on its strongest point. On the pricing test, it argued against its own retention finding. On the rollout, it gave Commercial no date to plan around.
- Tests and gates that can't settle the question: a 12-week pilot too short for a 6–10-week sales cycle, a pause trigger with no threshold, and a "90% sendable" bar that doesn't say which tickets count.
- Decision slides that ask for no decision, owners it made up, and brief wording left on the slides.
- Best setupGPT-6 Astra · ChatGPT90.3provisional
- Against the leader+17.1Opus 5.5 · Claude
Task by task
Its Combined score on each core task, across 19 outputs. Beside it, the mistakes a PM recorded in earlier reviews.
| Model | OverallAll core tasks | Define | Discover | Design | Experiment | Challenge | Operate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Build a roadmapProvisional | Develop product strategy | Write a PRD | Customer research call guideProvisional | Extract discovery insights | One-shot prototype | Activation & onboarding review | Experiment specificationProvisional | Analyse experiment results | 1000x an ideaProvisional | Challenge an idea | Write a stakeholder update | Make the launch call | ||
| GPT-6 AstraOpenAI | 90Partial | – | 81 | 95Best | 94 | 96Best | 70 | 99 | – | 95Best | – | 96 | 88Best | 89 |
| Opus 5.5Anthropic | 73#1 | 73 | 74 | 64 | 86 | 71 | 77 | 74 | 77 | 63 | 82 | 64 | 64 | 84 |
- Decision deferred5 of 19Write a stakeholder update, Make the launch call, Analyse experiment results
- Test or gate too weak4 of 19Develop product strategy, Make the launch call, Write a PRD
- Contradiction missed3 of 19Make the launch call, Analyse experiment results, Activation & onboarding review
- Hypothesis stated as fact2 of 19Challenge an idea, Activation & onboarding review
- Invented evidence1 of 19Write a PRD
- Numbers wrong1 of 19Analyse experiment results
- Artefact unusable1 of 19One-shot prototype
- Other1 of 19Make the launch call
What you’ll need to check
The fixes you’ll make most often, each with a real example.
Make the call
Lays out options or hedges where the brief asked for a decision the evidence can support.
Expand only where the evidence supports it; insufficient evidence means continued limited exposure.
Lays out options or hedges where the brief asked for a decision the evidence can support.
Make the call · targeted repairGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.
Read the outputTighten the test
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.
A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.
Tighten the test · quick editSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.
Read the outputSurface the contradiction
Smooths over conflicting evidence that changes the answer.
We also lose 29% before onboarding finishes.
Install → finished onboarding 71% → started trial 52%
Surface the contradiction · targeted repairAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.
Read the outputHow it compares, task by task
Its Combined score on each core task, next to every other ranked setup.
- Build a roadmap–
- Customer research call guide94.2leads
- Experiment specification–
- Develop product strategy80.6
- One-shot prototype69.6
- Write a stakeholder update88.3leads
- Make the launch call89.4
- 1000x an idea–
- Write a PRD95.3leads
- Extract discovery insights96.3leads
- Challenge an idea95.8
- Analyse experiment results95.1leads
- Activation & onboarding review98.8
GPT-6 Astra · ChatGPT Other ranked setups. Hover a dot for its score and runs.
| Task | Opus 5.5 · Claude | GPT-6 Astra · ChatGPT | GPT-6.1 Sol · API | GPT-6 Luna · API | Sonnet 5.5 · API | Gemini 3.5 Flash-Lite · Gemini |
|---|---|---|---|---|---|---|
| Build a roadmap | 72.5 (2 runs) | not run | 89.4 (2 runs) | not run | 92.9 (2 runs) | not run |
| Customer research call guide | 85.6 (2 runs) | 94.2 (1 runs) | 92.7 (2 runs) | 59.3 (2 runs) | 84.1 (2 runs) | not run |
| Experiment specification | 76.9 (2 runs) | not run | 87.5 (2 runs) | 76.0 (2 runs) | 86.5 (2 runs) | 23.1 (1 runs) |
| Develop product strategy | 73.9 (2 runs) | 80.6 (2 runs) | 87.2 (2 runs) | 89.4 (2 runs) | 83.5 (2 runs) | 57.7 (2 runs) |
| One-shot prototype | 76.8 (2 runs) | 69.6 (2 runs) | 88.7 (2 runs) | 84.5 (2 runs) | 92.9 (2 runs) | 49.4 (2 runs) |
| Write a stakeholder update | 64.4 (2 runs) | 88.3 (2 runs) | 81.2 (2 runs) | 82.2 (2 runs) | 67.0 (2 runs) | 50.2 (2 runs) |
| Make the launch call | 84.2 (2 runs) | 89.4 (2 runs) | 88.3 (2 runs) | 80.5 (2 runs) | 91.2 (2 runs) | 45.5 (2 runs) |
| 1000x an idea | 82.0 (2 runs) | not run | not run | 90.9 (1 runs) | not run | not run |
| Write a PRD | 64.0 (2 runs) | 95.3 (2 runs) | 68.9 (2 runs) | 66.7 (2 runs) | 75.6 (2 runs) | 41.7 (2 runs) |
| Extract discovery insights | 71.3 (2 runs) | 96.3 (2 runs) | 92.5 (2 runs) | 82.5 (2 runs) | 77.5 (2 runs) | 38.8 (2 runs) |
| Challenge an idea | 63.8 (2 runs) | 95.8 (2 runs) | 97.9 (2 runs) | 90.6 (2 runs) | 82.3 (2 runs) | 66.9 (2 runs) |
| Analyse experiment results | 63.1 (2 runs) | 95.1 (2 runs) | 90.5 (2 runs) | 90.5 (2 runs) | 79.7 (2 runs) | 52.5 (2 runs) |
| Activation & onboarding review | 73.8 (2 runs) | 98.8 (2 runs) | 100.0 (2 runs) | 90.0 (2 runs) | 57.5 (2 runs) | 55.0 (2 runs) |
Every tested harness
Same model, different setup. The harness often matters more than the launch post says.
| Rank | Harness | Combined score | Decision model | LLM judge | Evidence |
|---|---|---|---|---|---|
| Prov. | GPT-6 AstrawithChatGPTChatGPT 2026-09 | 90.7 | 92.1 |
By category
GPT-6 Astra · ChatGPT
Consistency
Typical time to output
Published runs
19 outputs, each graded blind on its task’s current checklist
| Task · case | Harness | Combined score |
|---|---|---|
| Activation & onboarding reviewAnalytics tool losing users at setup | ChatGPT | |
| Activation & onboarding reviewFitness app first week | ChatGPT | |
| Analyse experiment resultsThe underpowered onboarding test | ChatGPT | |
| Analyse experiment resultsConversion up, retention down | ChatGPT | |
| Challenge an ideaThe CEO's embedded-payments bet | ChatGPT | |
| Challenge an ideaAn AI SDR for small agencies | ChatGPT | |
| Customer research call guideWould freelancers pay to stop chasing invoices? | ChatGPT | |
| Develop product strategyHidden validation case | ChatGPT | |
| Develop product strategySupply or demand for a stalled marketplace | ChatGPT | |
| Extract discovery insightsHidden validation case | ChatGPT | |
| Extract discovery insightsEight calls with finance teams | ChatGPT | |
| Make the launch callGo/no-go for AI-drafted support replies | ChatGPT | |
| Make the launch callGo/no-go for AI substitutions, from the rollout data | ChatGPT | |
| One-shot prototypeExpense receipt capture | ChatGPT | |
| One-shot prototypeClinic appointment rebooking | ChatGPT | |
| Write a PRDMeeting summaries with action items | ChatGPT | |
| Write a PRDAI triage for support tickets | ChatGPT | |
| Write a stakeholder updateThe CEO asks why a metric dropped | ChatGPT | |
| Write a stakeholder updateA quarterly business review deck, in brand | ChatGPT |