By the XpertPress Editorial Team · Reviewed for technical accuracy · Last updated September 18, 2026
On September 3, 2026, OpenAI released GPT-6 Astra, the new model now powering ChatGPT, and immediately the internet split into two camps. One side pointed to a headline number — a 99.9% score on ARC-AGI-3, a benchmark specifically designed to resist AI progress — and declared that artificial general intelligence (AGI) had effectively arrived. The other side dug into the fine print and found a very different story: a 62.7% score under neutral testing conditions, an aggregate “intelligence index” that barely moved from the previous model, and a cybersecurity risk rating serious enough that OpenAI is gating parts of the model behind an approval program.
Both of those things are true at the same time. This article walks through exactly what changed in GPT-6 Astra, what the benchmark numbers really mean once you separate marketing from methodology, how it stacks up against GPT-5.6 Sol and Anthropic’s Claude models, and what any of this actually means if you run a website, a business, or just use ChatGPT day to day.
What Is GPT-6 Astra, Exactly?
GPT-6 Astra is OpenAI’s successor to GPT-5.6 Sol, and it’s the model now labeled “GPT-6 Astra” inside ChatGPT for Plus, Pro, Business, and Enterprise subscribers, with API access through OpenAI directly as well as Microsoft Azure and AWS Bedrock. OpenAI’s own launch materials call it “a significant jump in cyber capabilities” and its most capable model for computer use — meaning it can operate a browser or desktop environment on a user’s behalf to complete multi-step professional tasks.
Rather than being a single across-the-board upgrade, Astra is best understood as a specialist upgrade: it makes large, well-documented gains in a handful of specific areas — computer use, terminal/coding-agent work, and cybersecurity-adjacent reasoning — while showing flat or even mixed results on broader, aggregate measures of general intelligence.
Rollout and Access
- Launch date: September 3, 2026, starting with a limited set of organizations in OpenAI’s Trusted Access Program
- General availability: Rolling out through September 2026 to ChatGPT Plus, Pro, Business, and Enterprise plans
- API access: Available directly through OpenAI, plus Microsoft Azure and AWS Bedrock, with Zero Data Retention support for eligible customers
- A “Pro” variant is also available for Pro, Business, and Enterprise users
The AGI Claim: Where the 99.9% Number Actually Comes From
The single most-shared statistic from the Astra launch is a score of 99.9% on ARC-AGI-3 — a benchmark built by the ARC Prize Foundation specifically to test the kind of novel, abstract reasoning that pattern-matching language models have historically struggled with. Saturating a benchmark designed to stay ahead of AI capability is, on its face, a remarkable claim.
The catch, confirmed by multiple independent outlets that reviewed OpenAI’s own system card, is that the 99.9% figure came from a “Provider Adapter” harness — a testing setup that preserves reasoning state between actions in ways a standard evaluation doesn’t allow. When ARC Prize ran Astra through its neutral, standardized harness — the same conditions used to test every other model, including GPT-5.6 Sol — the score dropped to 62.7%.
That’s still a very large jump. Under that same neutral harness, GPT-5.6 Sol scored just 7.8%, meaning Astra’s real, apples-to-apples improvement is roughly an 8x gain — genuinely significant, just not the “the model basically solved general intelligence” headline that spread first.
So Is It AGI?
By the most common technical definition of AGI — a system that matches or exceeds human performance across essentially all cognitive tasks — no independent evaluator has concluded that GPT-6 Astra qualifies. Artificial Analysis’s own aggregate Intelligence Index, which blends dozens of benchmarks rather than spotlighting one, puts Astra at 61.2, essentially tied with its predecessor’s 60.9, and behind Anthropic’s Claude Fable 5.1 at 65.7. In other words: on the metric built specifically to avoid this kind of cherry-picking, GPT-6 Astra did not represent a generational leap.
What Astra does represent is a much sharper, more specialized jump in a narrower set of capabilities — and that distinction matters a lot for anyone deciding whether to actually pay for it.
Benchmark-by-Benchmark: Where Astra Actually Wins
Stripping out the headline-grabbing number, here’s where GPT-6 Astra’s gains are real, sourced from OpenAI’s own system card and cross-checked against independent labs including Artificial Analysis, Vellum, and ARC Prize:
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | What it measures |
|---|---|---|---|
| ARC-AGI-3 (standard harness) | 62.7% | 7.8% | Novel abstract reasoning |
| FrontierMath Tier 4 | 97.6% | — (unreported) | Advanced mathematics |
| OSWorld 2.0 (computer use) | 72.6%, ~47% faster | Baseline | Operating a real desktop/browser |
| ExploitBench (cyber) | 100% | 78.5% | Turning known vulnerabilities into working exploits |
| Terminal-Bench 4.0 | 56% | 37% | Command-line / dev-agent tasks |
| SRE-Bench (incident response) | 88.0% | 55.9% | First-attempt fix of production incidents |
| Humanity’s Last Exam (w/ tools) | 57.2% | — (unreported) | Expert-level cross-domain reasoning; Astra trails Claude Fable 5.1’s 65.0% |
| Artificial Analysis Intelligence Index | 61.2 | 60.9 | Broad aggregate score across ~dozens of evals |
Sources: OpenAI GPT-6 Astra system card (Sept 3, 2026); Artificial Analysis independent benchmarking; ARC Prize verified results; Vellum AI benchmark breakdown.
The pattern that emerges: Astra’s biggest, most defensible wins are in agentic and applied tasks — using a computer, working in a terminal, fixing production incidents, finding software exploits — rather than in raw, general reasoning. On several coding benchmarks (DeepSWE, Frontier Code), independent testers found Astra roughly tied with Claude Fable 5.1, Claude Opus 5, and even some competitor models rather than clearly ahead.
The Part Nobody’s Marketing Team Wanted to Lead With: Cybersecurity Risk
OpenAI itself classifies GPT-6 Astra as “Critical” for cybersecurity under its internal Preparedness Framework — the highest risk tier the company tracks. This isn’t a hypothetical concern: on ExploitBench, a benchmark that tests whether a model can turn a known software vulnerability into a working exploit, Astra scored a perfect 100%, up from 78.5% for GPT-5.6 Sol, and it did so using meaningfully fewer output tokens, i.e., more efficiently.
Because of this, OpenAI is shipping the model’s most sensitive exploit-creation capabilities gated behind a controlled access program (internally referred to as “Daybreak”), and the general rollout is staged more slowly than past model launches specifically because of this risk classification. Independent cybersecurity assessment firm Irregular reportedly corroborated the elevated capability level.
For everyday ChatGPT users this mostly shows up as invisible guardrails and refusal behavior. For security teams and IT decision-makers, it’s a meaningful signal: the same reasoning gains that make Astra good at fixing bugs and resolving infrastructure incidents also make it good at finding and weaponizing vulnerabilities.
Pricing: A 2.5x Jump in Cost per Token
GPT-6 Astra is noticeably more expensive to run than its predecessor through the API:
At $10 per million input tokens and $50 per million output tokens, Astra costs roughly 2.5x more per token than GPT-5.6 Sol’s $4/$20 pricing. OpenAI argues this is partly offset by efficiency: Astra tends to use fewer output tokens and fewer “turns” to complete a given agentic task — independent testing from Artificial Analysis found it used about a third of the tokens GPT-5.6 Sol needed for equivalent coding-agent tasks, and completed GDPval tasks in 24 turns on average versus 45 for Sol and 60 for Claude’s models.
Even accounting for that efficiency, real-world cost per completed task at maximum reasoning effort still runs meaningfully higher than Sol — Artificial Analysis measured about $7.09 per task for Astra versus roughly 15% less for Sol, though notably still 30–40% cheaper per task than Anthropic’s Claude Opus 5 or Claude Fable 5.1 at their own maximum settings for a comparable score.
GPT-6 Astra vs. the Competition: The Honest Summary
- vs. GPT-5.6 Sol (its own predecessor): Large, real gains in computer use, terminal/coding-agent tasks, and cybersecurity reasoning. Aggregate “general intelligence” essentially unchanged.
- vs. Claude Fable 5.1 (Anthropic, released Sept 1, 2026): Astra leads on computer use, terminal workflows, and long-context retrieval. Fable 5.1 leads on the broad Intelligence Index (65.7 vs 61.2) and on the Coding Agent Index (70 vs 67), and beats Astra on Humanity’s Last Exam with tools (65.0% vs 57.2%).
- vs. Claude Opus 5 (Anthropic): Astra’s ARC-AGI-3 standard-harness score (62.7%) beats Opus 5’s 30.2%, but Astra runs roughly 30% cheaper per task for a similar or better score on several agentic benchmarks.
The clearest, least disputed takeaway across independent reviewers: GPT-6 Astra is not a clean sweep. It’s a model with a few genuinely best-in-class specialties — especially anything involving operating a computer or a terminal autonomously — sitting inside an overall competitive landscape where no single lab has a decisive lead on general reasoning.
What This Actually Means If You Run a Website or a Business
For site owners, developers, and WordPress users specifically, the practical implications are more useful than the AGI debate itself:
1. Agentic, “do it for me” AI is getting genuinely more reliable
The OSWorld 2.0 and terminal-benchmark gains are the numbers that matter most for anyone using AI to automate real work — content publishing, QA testing, repetitive plugin configuration, or multi-step admin tasks. A model that completes tasks in fewer turns and with less hand-holding is the trend to watch, regardless of whether it’s technically “AGI.”
2. Treat headline AI benchmark claims skeptically by default
The gap between Astra’s 99.9% and 62.7% ARC-AGI-3 scores is a useful case study for evaluating any AI vendor’s marketing claims going forward: always ask what harness, what tooling, and what conditions produced the number before treating it as representative.
3. Cost-conscious teams should benchmark their own workloads
A 2.5x per-token price increase is significant at scale. If you’re building AI-powered features into a WordPress plugin or e-commerce workflow, the efficiency gains (fewer tokens, fewer turns) may or may not offset the higher per-token price depending on your specific task — this is worth testing directly rather than assuming.
Frequently Asked Questions
Is GPT-6 Astra actually AGI?
No independent evaluator has concluded that GPT-6 Astra meets standard definitions of artificial general intelligence. Its widely cited 99.9% ARC-AGI-3 score came from a non-standard testing harness; under neutral testing conditions it scored 62.7%, and its broad aggregate intelligence score was essentially flat versus its predecessor.
What is GPT-6 Astra’s real ARC-AGI-3 score?
62.7% under ARC Prize’s standard, neutral harness — still roughly 8x higher than GPT-5.6 Sol’s 7.8% on the same test, but well below the 99.9% figure OpenAI highlighted using its own provider-specific testing setup.
Why is GPT-6 Astra rated “Critical” for cybersecurity?
Under OpenAI’s Preparedness Framework, Astra crossed the “Critical” threshold for cyber capability after scoring 100% on ExploitBench, a benchmark measuring the ability to turn known vulnerabilities into working exploits. As a result, its most sensitive capabilities are gated behind a controlled access program.
How much does GPT-6 Astra cost compared to GPT-5.6 Sol?
Astra’s API pricing is $10 per million input tokens and $50 per million output tokens, versus $4/$20 for GPT-5.6 Sol — roughly 2.5x more expensive per token, partially offset by using fewer tokens per completed task.
Is GPT-6 Astra better than Claude?
It depends on the task. Astra leads on computer-use and terminal/agentic benchmarks. Anthropic’s Claude Fable 5.1 leads on the broad Artificial Analysis Intelligence Index and on expert-level reasoning with tools. Neither model has a decisive lead across the board.
Sources
- OpenAI, “GPT-6 Astra: A new generation of intelligence” — official launch page and system card, September 3, 2026
- Artificial Analysis, “Benchmarking GPT-6 Astra” — independent benchmark suite, September 2026
- ARC Prize Foundation — verified ARC-AGI-3 standard-harness results
- Vellum AI, “GPT-6 Astra Benchmarks Explained”
- DataCamp, “GPT-6 Astra: Features, Benchmarks, and Pricing”
- MindStudio, “GPT-6 Astra Benchmarks: Is It Really Better Than Fable 5.1?”
- Generative AI (pub), “GPT-6 Astra Benchmarks: 99.9% or 62.7%? The Full ARC-AGI-3 Story”
Editorial note: Benchmark scores and pricing for frontier AI models change frequently. This article reflects publicly available data as of September 18, 2026, and will be updated if OpenAI or independent evaluators publish revised figures.
