Ratings / OpenAI

GPT-6 Astra

OpenAI · gpt-6-astra · effort medium (default)

GPT-6 AstrawithChatGPTProvisional
19 runs · ±4.4 · Moderate evidence

What you can hand it, and what you’ll still have to catch

Dropped 25 Sept 2026, against whatever leads today.

Handles most PM work well, but sometimes won't commit.

Fifteen of Astra's 18 outputs were usable with at most a quick edit, and six were ready to use as written. It keeps what the evidence shows apart from what it suggests, better than any setup so far. When it falls short, it hedges where a decision was needed, or it proposes a test that couldn't settle the question.

What you can hand it

  • Discovery synthesis. Both cases were rated Ready to use, with confidence matched to the evidence.
  • Strategy memos and challenges to a plan. It separates evidence from inference and shows its working.
  • Go/no-go calls against agreed criteria. It applies each gate as written, and doesn't let human review excuse a failure.
  • PRDs where the safety constraint is the hard part.

What you’ll still have to catch

  • Hedging on its strongest point. On the pricing test, it argued against its own retention finding. On the rollout, it gave Commercial no date to plan around.
  • Tests and gates that can't settle the question: a 12-week pilot too short for a 6–10-week sales cycle, a pause trigger with no threshold, and a "90% sendable" bar that doesn't say which tickets count.
  • Decision slides that ask for no decision, owners it made up, and brief wording left on the slides.
  • Best setupGPT-6 Astra · ChatGPT90.3provisional
  • Against the leader+17.1Opus 5.5 · Claude

Task by task

Its Combined score on each core task, across 19 outputs. Beside it, the mistakes a PM recorded in earlier reviews.

GPT-6 Astra's Combined score on each task, beside the current leader.
ModelOverallAll core tasksDefineDiscoverDesignExperimentChallengeOperate
Build a roadmapProvisionalDevelop product strategyWrite a PRDCustomer research call guideProvisionalExtract discovery insightsOne-shot prototypeActivation & onboarding reviewExperiment specificationProvisionalAnalyse experiment results1000x an ideaProvisionalChallenge an ideaWrite a stakeholder updateMake the launch call
GPT-6 AstraOpenAI90Partial–8195Best9496Best7099–95Best–9688Best89
Opus 5.5Anthropic73#173746486717774776382646484
Combined scoreEach cell is the model’s best setup on that task. Provisional tasks are still being calibrated.

What you’ll need to check

The fixes you’ll make most often, each with a real example.

Make the call

Lays out options or hedges where the brief asked for a decision the evidence can support.

GPT-6 AstrawithChatGPT on Go/no-go for AI substitutions, from the rollout data
It wrote
Expand only where the evidence supports it; insufficient evidence means continued limited exposure.
What was missing

Lays out options or hedges where the brief asked for a decision the evidence can support.

Make the call · targeted repairGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.

Read the output

Tighten the test

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

GPT-6 AstrawithChatGPT on AI triage for support tickets
It wrote
at least 90% rated sendable unchanged or with minor stylistic edits by support reviewers.
What was missing

A success measure, gate, kill criterion or test that couldn’t settle the question it’s there to settle.

Tighten the test · quick editSay which tickets count towards the 90%, and how unavailable drafts and abstentions are reported, so the gate can't be met by drafting only the easy tickets. Say how the confidence intervals affect passing.

Read the output

Surface the contradiction

Smooths over conflicting evidence that changes the answer.

GPT-6 AstrawithChatGPT on Fitness app first week
It wrote
We also lose 29% before onboarding finishes.
The source (Funnel)

Install → finished onboarding 71% → started trial 52%

Surface the contradiction · targeted repairAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.

Read the output

How it compares, task by task

Its Combined score on each core task, next to every other ranked setup.

Compare any setup
  1. Build a roadmap–
  2. Customer research call guide94.2leads
  3. Experiment specification–
  4. Develop product strategy80.6
  5. One-shot prototype69.6
  6. Write a stakeholder update88.3leads
  7. Make the launch call89.4
  8. 1000x an idea–
  9. Write a PRD95.3leads
  10. Extract discovery insights96.3leads
  11. Challenge an idea95.8
  12. Analyse experiment results95.1leads
  13. Activation & onboarding review98.8

GPT-6 Astra · ChatGPT Other ranked setups. Hover a dot for its score and runs.

Combined score by task
TaskOpus 5.5 · ClaudeGPT-6 Astra · ChatGPTGPT-6.1 Sol · APIGPT-6 Luna · APISonnet 5.5 · APIGemini 3.5 Flash-Lite · Gemini
Build a roadmap72.5 (2 runs)not run89.4 (2 runs)not run92.9 (2 runs)not run
Customer research call guide85.6 (2 runs)94.2 (1 runs)92.7 (2 runs)59.3 (2 runs)84.1 (2 runs)not run
Experiment specification76.9 (2 runs)not run87.5 (2 runs)76.0 (2 runs)86.5 (2 runs)23.1 (1 runs)
Develop product strategy73.9 (2 runs)80.6 (2 runs)87.2 (2 runs)89.4 (2 runs)83.5 (2 runs)57.7 (2 runs)
One-shot prototype76.8 (2 runs)69.6 (2 runs)88.7 (2 runs)84.5 (2 runs)92.9 (2 runs)49.4 (2 runs)
Write a stakeholder update64.4 (2 runs)88.3 (2 runs)81.2 (2 runs)82.2 (2 runs)67.0 (2 runs)50.2 (2 runs)
Make the launch call84.2 (2 runs)89.4 (2 runs)88.3 (2 runs)80.5 (2 runs)91.2 (2 runs)45.5 (2 runs)
1000x an idea82.0 (2 runs)not runnot run90.9 (1 runs)not runnot run
Write a PRD64.0 (2 runs)95.3 (2 runs)68.9 (2 runs)66.7 (2 runs)75.6 (2 runs)41.7 (2 runs)
Extract discovery insights71.3 (2 runs)96.3 (2 runs)92.5 (2 runs)82.5 (2 runs)77.5 (2 runs)38.8 (2 runs)
Challenge an idea63.8 (2 runs)95.8 (2 runs)97.9 (2 runs)90.6 (2 runs)82.3 (2 runs)66.9 (2 runs)
Analyse experiment results63.1 (2 runs)95.1 (2 runs)90.5 (2 runs)90.5 (2 runs)79.7 (2 runs)52.5 (2 runs)
Activation & onboarding review73.8 (2 runs)98.8 (2 runs)100.0 (2 runs)90.0 (2 runs)57.5 (2 runs)55.0 (2 runs)

Every tested harness

Same model, different setup. The harness often matters more than the launch post says.

RankHarnessCombined scoreDecision modelLLM judgeEvidence
Prov.GPT-6 AstrawithChatGPTChatGPT 2026-0990.792.1

Strengths

    Leads: Customer research call guide, Write a stakeholder update, Write a PRD, Extract discovery insights, Analyse experiment results

    Recurring failures

      Bottom half on: Develop product strategy, One-shot prototype

      By category

      GPT-6 Astra · ChatGPT

      Define88
      Discover95
      Design84
      Experiment95
      Challenge96
      Operate89

      Consistency

      1.5 min

      Typical time to output

      Published runs

      19 outputs, each graded blind on its task’s current checklist

      Task · caseHarnessCombined score
      Activation & onboarding reviewAnalytics tool losing users at setupChatGPT
      Activation & onboarding reviewFitness app first weekChatGPT
      Analyse experiment resultsThe underpowered onboarding testChatGPT
      Analyse experiment resultsConversion up, retention downChatGPT
      Challenge an ideaThe CEO's embedded-payments betChatGPT
      Challenge an ideaAn AI SDR for small agenciesChatGPT
      Customer research call guideWould freelancers pay to stop chasing invoices?ChatGPT
      Develop product strategyHidden validation caseChatGPT
      Develop product strategySupply or demand for a stalled marketplaceChatGPT
      Extract discovery insightsHidden validation caseChatGPT
      Extract discovery insightsEight calls with finance teamsChatGPT
      Make the launch callGo/no-go for AI-drafted support repliesChatGPT
      Make the launch callGo/no-go for AI substitutions, from the rollout dataChatGPT
      One-shot prototypeExpense receipt captureChatGPT
      One-shot prototypeClinic appointment rebookingChatGPT
      Write a PRDMeeting summaries with action itemsChatGPT
      Write a PRDAI triage for support ticketsChatGPT
      Write a stakeholder updateThe CEO asks why a metric droppedChatGPT
      Write a stakeholder updateA quarterly business review deck, in brandChatGPT