Skip to content

Tech Innovations

GPT-6 Astra Is Live: What OpenAI Actually Shipped, What the Benchmarks Mean, and Whether the “AGI Era” Started

OpenAI launched GPT-6 Astra on September 3, 2026. Paid ChatGPT plans, the API, Azure, and AWS Bedrock are getting it in stages. Company president Greg Brockman used the phrase “AGI era.” That sentence is now everywhere.

By Benjamin Thomas 6 min read
GPT-6 Astra Is Live: What OpenAI Actually Shipped, What the Benchmarks Mean, and Whether the “AGI Era” Started

OpenAI launched GPT-6 Astra on September 3, 2026. Paid ChatGPT plans, the API, Azure, and AWS Bedrock are getting it in stages. Company president Greg Brockman used the phrase “AGI era.” That sentence is now everywhere.

This article separates three things that keep getting mashed together:

  1. What Astra can do in products you can actually use
  2. What the published scores measure — and what they do not
  3. Whether calling this AGI is a definition, a forecast, or a press line

Short version: Astra is real, it is rolling out, it is expensive, and it is strongest on computer use, coding agents, math, and long tasks. The 99.9% ARC-AGI-3 number is not the same test as the 62.7% number. Those two scores are the most useful fact of the launch.

What is GPT-6 Astra?

GPT-6 Astra is OpenAI’s new flagship after GPT-5.6 Sol. Not a small point release. OpenAI positions it for end-to-end work: using a computer, browsing, writing software, professional workflows, science, and cybersecurity.

API model ID: gpt-6-astra

Spec Value
Context window 1,050,000 tokens
Max input 922,000 tokens
Max output 128,000 tokens
Knowledge cutoff April 30, 2026
Input Text and images
Output Text
Audio / video Not on this model
Fine-tuning Not supported

Pro, Business, and Enterprise ChatGPT plans also get GPT-6 Astra Pro, which burns more weekly quota.

Is GPT-6 Astra live right now?

Yes. Staged rollout.

  • September 3, 2026: announced; limited orgs via Daybreak / early access
  • Following days: ChatGPT Plus, Pro, Business, Enterprise
  • Same window: OpenAI API, Microsoft Azure Foundry, AWS Bedrock

Enterprise workspaces often ship with Astra off by default. An admin has to turn it on. Usage sits inside existing plan allowances; extra usage is credits. Astra eats quota faster than Sol. Free ChatGPT does not get the flagship.

GPT-6 Astra pricing (API)

Token type Price per 1M tokens
Input $10.00
Cached input $1.00
Cache writes $12.50
Output $50.00

Rules that change the bill:

  • Prompts over 272K input tokens bill the whole request at 2× input/cache and 1.5× output
  • Batch / Flex: 50% of standard
  • Fast mode: 2× price for about 2× speed
  • Computer use and search add per-tool-call fees

Versus GPT-5.6 Sol at $4 / $20, Astra is about 2.5× on tokens. Claude Fable 5.1 is the same $10 / $50 sticker. Claude Opus 5 is $5 / $25. For long agent runs, output tokens dominate. Cache the system prompt.

What Astra is actually good at

Computer and browser use

OpenAI’s product story. OSWorld 2.0: 72.6% in about 40 minutes per task, versus Sol’s 65.7% in about 75 minutes. ScreenSpot-Pro: 92.7% vs Sol’s 76.9%.

Translation: forms, CRM, calendars, job/apartment search, real Excel/Power BI work, installs, admin panels. Demos include KiCad layouts, Blender-to-Unreal scenes, and ChatGPT Sites. 72.6% is not “better than a competent human at every desktop task.” It is faster than last year’s agent and still needs supervision.

Software engineering

DeepSWE v1.1: 74.1% (Sol 72.7%). Terminal-Bench 4.0: 57.9% (Sol 37.3%, Fable 5.1 55.8%). The Terminal-Bench jump is the one that matters for real repos. Codex is getting a harness update; OpenAI claims 1.9× faster Mind2Web with the new computer-use path.

Math and science

FrontierMath Tier 4: about 98%. OpenAI says Astra helped tighten two number-theory bounds (prime pairs at most 186 apart; a large-gap term unchanged for 80+ years), with posted proofs and human verification. That is research help, not “the model is a number theorist.”

GPQA Diamond: 96.0%. Humanity’s Last Exam with tools: 57.2%, behind Fable 5.1 at 65.0%. Not uniformly first.

Cybersecurity

ExploitBench: 100%. OpenAI rates it Critical: unknown flaws plus working attacks without a human pointing at the hole. Consumer/API surfaces refuse advanced offensive work. Daybreak is the defender lane.

Documents and “professional work”

Stronger template-following for decks, workbooks, and sites. Agents’ Last Exam: 59.3%. AutomationBench: 41.4% vs Sol’s 18.1%. Early testers split: spatial/agentic work got praise; several said the prose is worse than Sol. One index reported roughly an 80 Elo drop on economically valuable professional work. Test your own docs before switching the default.

The ARC-AGI-3 number, without the marketing

ARC Prize ran Astra two ways:

Condition Score Approx. eval cost
Standard harness (notes only) 62.7% at max ~$26,000
Provider Adapter (hidden reasoning + compaction, closer to ChatGPT/Codex) 99.9% at high ~$19,000

OpenAI’s launch chart used 99.9% against Sol’s 7.8% and Opus 5’s ~30% from the standard harness. Different exams.

ARC Prize is not claiming AGI. Mike Knoop: they “lack evidence to call this AGI yet.” 62.7% is still a large jump from 7.8%. 99.9% is what the product stack can do when private thought survives between clicks. Both are real. Only one is a clean model-vs-model number.

Did the AGI era just start?

OpenAI’s working bar is a system that outperforms humans at most economically valuable work. Brockman said people will later point at “this time and this model.”

Reasons not to treat that as settled:

  1. Most jobs are not a leaderboard.
  2. Harnesses move scores more than headlines admit.
  3. The model is uneven (HLE, some writing).
  4. It is a hosted product with quotas, a April 30, 2026 cutoff, and refused tools.
  5. The benchmark owner declined the AGI label.

Better sentence: Astra is the first OpenAI model the company will sell as the start of an AGI era — agentic computer use, hard-eval saturation, junior-scientist assists. Chollet’s narrower point is more useful: symbolic world-modeling that used to live in fancy harnesses is moving into the model. That is a shift. It is not “most jobs are optional.”

GPT-6 Astra vs Sol vs Fable 5.1

Job Default pick (Sept 2026) Why
Desktop / browser agents Astra OSWorld + ScreenSpot-Pro
Long terminal / repo agents Astra Terminal-Bench 4.0
Cheap “good enough” volume Sol ($4/$20) ~2.5× cheaper
HLE + some writing Fable 5.1 65% vs 57.2%; voice reports
FrontierMath Astra Near saturation
Everyday ChatGPT prose Test both Mixed early reports

Same $10/$50 list price as Fable 5.1. A/B the task.

How to try it

ChatGPT: paid plan → model picker → Astra / Astra Pro. Ask the admin if you are on Enterprise.

API: gpt-6-astra on Chat Completions, Responses, Batch. No Realtime on this model. Image in, text out.

Cloud: Azure Foundry, AWS Bedrock.

Cost control: cache the system prompt, leave effort off max unless medium fails, stay under 272K input, use Batch overnight, measure dollars per finished task.

Safety that affects real use

OpenAI says Astra stayed in authorized scope in 0% of a new overreach eval vs 48% for unsafeguarded Sol. Internal computer-use misalignment 2.4% vs 22%. They also say written reasoning is harder to monitor than Sol’s traces, and Critical cyber is why offensive tools are locked. Treat a desktop agent like a privileged intern: scoped accounts, logs, human approve on money, mail, and deploys.

FAQ

When did it launch? September 3, 2026.

Is it AGI? OpenAI says the era started. ARC Prize says they lack evidence. Mixed benches. Don’t outsource the definition to a launch post.

Free tier? No.

Context window? 1.05M tokens, 128K max out, cutoff April 30, 2026.

Why 99.9% and 62.7%? Different harnesses. Compare 62.7% to other models’ standard-harness scores.

Better writer than Sol? Not clearly. Test.

Replace Fable 5.1? No single winner.

Bottom line

GPT-6 Astra is live. Serious computer-use and agent model, near-saturated math evals, Critical cyber rating behind locks, price that assumes you only call it when the cheaper model fails.

The AGI-era headline is OpenAI’s. The evidence: agents got faster, fair ARC-AGI-3 went from single digits to the low sixties, an adapter score hit 99.9%, two number-theory bounds moved with checked proofs. That is a lot. It is not a labor-market event by itself.

Use the model. Keep the slogan in quotes.


Discover more from CortexHub

Subscribe to get the latest posts sent to your email.

Newsletter

The CortexHub Daily

The tech stories that matter — startups, AI, and the products people actually use. One email.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Redirecting to checkout…

Discover more from CortexHub

Subscribe now to keep reading and get access to the full archive.

Continue reading