Tech Innovations
GPT-6 Astra Is Live: What OpenAI Actually Shipped, What the Benchmarks Mean, and Whether the “AGI Era” Started
OpenAI launched GPT-6 Astra on September 3, 2026. Paid ChatGPT plans, the API, Azure, and AWS Bedrock are getting it in stages. Company president Greg Brockman used the phrase “AGI era.” That sentence is now everywhere.
OpenAI launched GPT-6 Astra on September 3, 2026. Paid ChatGPT plans, the API, Azure, and AWS Bedrock are getting it in stages. Company president Greg Brockman used the phrase “AGI era.” That sentence is now everywhere.
This article separates three things that keep getting mashed together:
- What Astra can do in products you can actually use
- What the published scores measure — and what they do not
- Whether calling this AGI is a definition, a forecast, or a press line
Short version: Astra is real, it is rolling out, it is expensive, and it is strongest on computer use, coding agents, math, and long tasks. The 99.9% ARC-AGI-3 number is not the same test as the 62.7% number. Those two scores are the most useful fact of the launch.
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s new flagship after GPT-5.6 Sol. Not a small point release. OpenAI positions it for end-to-end work: using a computer, browsing, writing software, professional workflows, science, and cybersecurity.
API model ID: gpt-6-astra
| Spec | Value |
|---|---|
| Context window | 1,050,000 tokens |
| Max input | 922,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Input | Text and images |
| Output | Text |
| Audio / video | Not on this model |
| Fine-tuning | Not supported |
Pro, Business, and Enterprise ChatGPT plans also get GPT-6 Astra Pro, which burns more weekly quota.
Is GPT-6 Astra live right now?
Yes. Staged rollout.
- September 3, 2026: announced; limited orgs via Daybreak / early access
- Following days: ChatGPT Plus, Pro, Business, Enterprise
- Same window: OpenAI API, Microsoft Azure Foundry, AWS Bedrock
Enterprise workspaces often ship with Astra off by default. An admin has to turn it on. Usage sits inside existing plan allowances; extra usage is credits. Astra eats quota faster than Sol. Free ChatGPT does not get the flagship.
GPT-6 Astra pricing (API)
| Token type | Price per 1M tokens |
|---|---|
| Input | $10.00 |
| Cached input | $1.00 |
| Cache writes | $12.50 |
| Output | $50.00 |
Rules that change the bill:
- Prompts over 272K input tokens bill the whole request at 2× input/cache and 1.5× output
- Batch / Flex: 50% of standard
- Fast mode: 2× price for about 2× speed
- Computer use and search add per-tool-call fees
Versus GPT-5.6 Sol at $4 / $20, Astra is about 2.5× on tokens. Claude Fable 5.1 is the same $10 / $50 sticker. Claude Opus 5 is $5 / $25. For long agent runs, output tokens dominate. Cache the system prompt.
What Astra is actually good at
Computer and browser use
OpenAI’s product story. OSWorld 2.0: 72.6% in about 40 minutes per task, versus Sol’s 65.7% in about 75 minutes. ScreenSpot-Pro: 92.7% vs Sol’s 76.9%.
Translation: forms, CRM, calendars, job/apartment search, real Excel/Power BI work, installs, admin panels. Demos include KiCad layouts, Blender-to-Unreal scenes, and ChatGPT Sites. 72.6% is not “better than a competent human at every desktop task.” It is faster than last year’s agent and still needs supervision.
Software engineering
DeepSWE v1.1: 74.1% (Sol 72.7%). Terminal-Bench 4.0: 57.9% (Sol 37.3%, Fable 5.1 55.8%). The Terminal-Bench jump is the one that matters for real repos. Codex is getting a harness update; OpenAI claims 1.9× faster Mind2Web with the new computer-use path.
Math and science
FrontierMath Tier 4: about 98%. OpenAI says Astra helped tighten two number-theory bounds (prime pairs at most 186 apart; a large-gap term unchanged for 80+ years), with posted proofs and human verification. That is research help, not “the model is a number theorist.”
GPQA Diamond: 96.0%. Humanity’s Last Exam with tools: 57.2%, behind Fable 5.1 at 65.0%. Not uniformly first.
Cybersecurity
ExploitBench: 100%. OpenAI rates it Critical: unknown flaws plus working attacks without a human pointing at the hole. Consumer/API surfaces refuse advanced offensive work. Daybreak is the defender lane.
Documents and “professional work”
Stronger template-following for decks, workbooks, and sites. Agents’ Last Exam: 59.3%. AutomationBench: 41.4% vs Sol’s 18.1%. Early testers split: spatial/agentic work got praise; several said the prose is worse than Sol. One index reported roughly an 80 Elo drop on economically valuable professional work. Test your own docs before switching the default.
The ARC-AGI-3 number, without the marketing
ARC Prize ran Astra two ways:
| Condition | Score | Approx. eval cost |
|---|---|---|
| Standard harness (notes only) | 62.7% at max | ~$26,000 |
| Provider Adapter (hidden reasoning + compaction, closer to ChatGPT/Codex) | 99.9% at high | ~$19,000 |
OpenAI’s launch chart used 99.9% against Sol’s 7.8% and Opus 5’s ~30% from the standard harness. Different exams.
ARC Prize is not claiming AGI. Mike Knoop: they “lack evidence to call this AGI yet.” 62.7% is still a large jump from 7.8%. 99.9% is what the product stack can do when private thought survives between clicks. Both are real. Only one is a clean model-vs-model number.
Did the AGI era just start?
OpenAI’s working bar is a system that outperforms humans at most economically valuable work. Brockman said people will later point at “this time and this model.”
Reasons not to treat that as settled:
- Most jobs are not a leaderboard.
- Harnesses move scores more than headlines admit.
- The model is uneven (HLE, some writing).
- It is a hosted product with quotas, a April 30, 2026 cutoff, and refused tools.
- The benchmark owner declined the AGI label.
Better sentence: Astra is the first OpenAI model the company will sell as the start of an AGI era — agentic computer use, hard-eval saturation, junior-scientist assists. Chollet’s narrower point is more useful: symbolic world-modeling that used to live in fancy harnesses is moving into the model. That is a shift. It is not “most jobs are optional.”
GPT-6 Astra vs Sol vs Fable 5.1
| Job | Default pick (Sept 2026) | Why |
|---|---|---|
| Desktop / browser agents | Astra | OSWorld + ScreenSpot-Pro |
| Long terminal / repo agents | Astra | Terminal-Bench 4.0 |
| Cheap “good enough” volume | Sol ($4/$20) | ~2.5× cheaper |
| HLE + some writing | Fable 5.1 | 65% vs 57.2%; voice reports |
| FrontierMath | Astra | Near saturation |
| Everyday ChatGPT prose | Test both | Mixed early reports |
Same $10/$50 list price as Fable 5.1. A/B the task.
How to try it
ChatGPT: paid plan → model picker → Astra / Astra Pro. Ask the admin if you are on Enterprise.
API: gpt-6-astra on Chat Completions, Responses, Batch. No Realtime on this model. Image in, text out.
Cloud: Azure Foundry, AWS Bedrock.
Cost control: cache the system prompt, leave effort off max unless medium fails, stay under 272K input, use Batch overnight, measure dollars per finished task.
Safety that affects real use
OpenAI says Astra stayed in authorized scope in 0% of a new overreach eval vs 48% for unsafeguarded Sol. Internal computer-use misalignment 2.4% vs 22%. They also say written reasoning is harder to monitor than Sol’s traces, and Critical cyber is why offensive tools are locked. Treat a desktop agent like a privileged intern: scoped accounts, logs, human approve on money, mail, and deploys.
FAQ
When did it launch? September 3, 2026.
Is it AGI? OpenAI says the era started. ARC Prize says they lack evidence. Mixed benches. Don’t outsource the definition to a launch post.
Free tier? No.
Context window? 1.05M tokens, 128K max out, cutoff April 30, 2026.
Why 99.9% and 62.7%? Different harnesses. Compare 62.7% to other models’ standard-harness scores.
Better writer than Sol? Not clearly. Test.
Replace Fable 5.1? No single winner.
Bottom line
GPT-6 Astra is live. Serious computer-use and agent model, near-saturated math evals, Critical cyber rating behind locks, price that assumes you only call it when the cheaper model fails.
The AGI-era headline is OpenAI’s. The evidence: agents got faster, fair ARC-AGI-3 went from single digits to the low sixties, an adapter score hit 99.9%, two number-theory bounds moved with checked proofs. That is a lot. It is not a labor-market event by itself.
Use the model. Keep the slogan in quotes.
Discover more from CortexHub
Subscribe to get the latest posts sent to your email.
