GPT-6 Astra: What OpenAI Claims vs What Independent Tests Show
OpenAI calls GPT-6 Astra its most intelligent and aligned model ever. Independent benchmarks tell a narrower story — real agentic gains, a flat general-intelligence score, and a new Critical cybersecurity risk rating.
OpenAI released GPT-6 Astra on September 3, 2026, calling it "the world's most intelligent and aligned model." Within 48 hours, independent evaluators had a different headline: strong, confirmed gains in a few specific areas, and no measurable jump in general intelligence.
Both things are true. The gap between them is the actual story.
What Astra is
Astra is a single reasoning model (gpt-6-astra in the API) with five effort settings — low, medium, high, xhigh, max. It has a 1.05M token context window, a knowledge cutoff of April 30, 2026, and costs $10 per million input tokens / $50 per million output tokens — roughly 2.5x the price of its predecessor, GPT-5.6 Sol. It's available now in ChatGPT Plus/Pro/Business/Enterprise, the OpenAI API, Microsoft Foundry, and AWS Bedrock.
The pitch is computer use: filling forms, navigating CRMs, running terminal workflows, doing agentic multi-step tasks with less supervision than any prior OpenAI model.
Where the numbers genuinely hold up
On OpenAI's own tables, Astra saturates FrontierMath Tier 4 (97.6%), ARC-AGI-3 (99.9%), and ExploitBench (100%). Those are real, but each comes with a footnote worth knowing:
- ARC-AGI-3's 99.9% was scored using OpenAI's own "provider adapter" harness. On the standard, harness-neutral evaluation, ARC Prize measured 62.7% — still frontier-level, but nowhere near saturation. Greg Kamradt's team, which built the benchmark, was explicit that beating it isn't evidence of AGI.
- FrontierMath saturation doesn't mean math is "solved." A separate, harder benchmark of unsolved Erdős problems had Astra solve 2 of 68 on its first pass, rising to 5 with repeated attempts that reportedly cost over $220,000 in compute.
Where the gains are independently confirmed, not just vendor-reported: computer use and coding efficiency. Artificial Analysis's Coding Agent Index puts Astra roughly level with Claude Opus 5 and Fable 5, but at a fraction of the token cost — matching Fable 5's score at less than half the price per task. On OSWorld 2.0, Astra scores higher than Sol in about 47% less time per task. That's a real, repeatable efficiency win.
Where the "most intelligent" claim doesn't survive contact
On Artificial Analysis's broad Intelligence Index — the closest thing to a neutral general-capability score — Astra lands at 61, statistically identical to its predecessor GPT-5.6 Sol, and behind Claude Fable 5.1 at roughly 66. At 2.5x the token price, that's about 75% more cost per task for a flat score.
Artificial Analysis also recorded a drop of roughly 80 Elo points on GDPval-AA v2 (a benchmark built from OpenAI's own dataset of economically valuable, real-world tasks across 44 occupations), plus smaller regressions in customer support and long-context reasoning tasks.
So the accurate read is: big, real wins in a narrow band — agentic computer use, terminal workflows, coding cost-efficiency — with no aggregate intelligence gain and a meaningful cost increase everywhere else.
The part that actually matters more than any benchmark: cybersecurity
This is the detail most coverage buries. OpenAI itself states Astra is the first model to meet its Critical threshold for cybersecurity capability under its Preparedness Framework — meaning that, given the right tools and access, it can find previously unknown vulnerabilities and build working exploit chains largely without human guidance.
The numbers behind that claim:
| Benchmark | Astra | GPT-5.6 Sol |
|---|---|---|
| ExploitBench | 100.0% | 78.5% |
| ExploitGym | 42.4% | 30.3% |
| SRE-Bench (binary reverse engineering) | 88.0% | 55.9% |
During internal testing on a fresh set of vulnerabilities from the prior three months (to rule out benchmark contamination), Astra discovered and used two previously unknown zero-day exploits. OpenAI disclosed both to the affected maintainers rather than publishing them.
Because of this, the publicly deployed version deliberately refuses advanced offensive tasks — proof-of-concept exploit generation, for instance — with less-restricted access gated behind OpenAI's "Daybreak" program for vetted defenders. It's a genuinely unusual thing to see a lab say plainly about its own model: this is more dangerous than what came before, and we've hobbled it on purpose.
Alignment: the good and the caveat worth flagging
OpenAI reports real improvements here too — Astra circumvented safety review denials 0% of the time in testing, versus meaningfully higher rates for Sol, and made fewer inaccurate claims about its own capabilities. That's a legitimate, testable improvement in agent reliability.
The caveat, from OpenAI's own system card: Astra's chain-of-thought reasoning is harder to monitor than its predecessor's, in part because it solves problems in fewer, more compressed steps. OpenAI flags this as a research priority rather than a solved problem. It's worth knowing if you're planning to give this model long-running, low-supervision tasks.
What this means for a build decision
If your use case is long-horizon agentic work — computer use, terminal-heavy coding, tasks that used to require constant babysitting — the efficiency and reliability gains are real and worth testing against your own workload. If your use case is general reasoning, writing, or analysis where token cost matters and you're not doing agentic loops, the case for Astra over cheaper alternatives is weaker on the current independent numbers.
Either way, evaluate on your own tasks before committing budget — vendor benchmarks and neutral benchmarks are measuring different things, and the 2.5x price jump only makes sense if you're actually using the agentic capability it's priced for. This is the kind of model-selection question we work through with clients directly inside our AI Strategy & Consulting engagements, and it feeds into how we scope AI Product Development work where the model choice determines the cost structure of the whole product.
If you're weighing GPT-6 Astra against other frontier models for a specific build, get in touch — we can run the comparison against your actual workload rather than a leaderboard.
Topics
Build this into a product
Talk to NTechnologies about applying these ideas in your stack.