OpenAI released GPT-6 Astra on Thursday and its president said the company had entered the era of artificial general intelligence. The company led its launch case with a 99.9% score on ARC-AGI-3, a test of how agents learn unfamiliar interactive environments. The organization that built the benchmark produced a 62.7% result under its provider-neutral setup and said it was not claiming Astra is AGI.
What Changed
- OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman closed the press briefing by saying "Welcome to the AGI era."
- OpenAI led its launch case with a 99.9% ARC-AGI-3 score. ARC Prize, which built the benchmark, scored the same model at 62.7% on its provider-neutral Standard harness and said it is not claiming Astra is AGI.
- Artificial Analysis put Astra at 61 on its Intelligence Index v4.1.1, level with GPT-5.6 Sol and about five points behind Claude Fable 5.1. OpenAI's own table shows Astra at 57.2% on Humanity's Last Exam with tools, below Sol's 65.0%.
- Standard API pricing is $10 per million input tokens and $50 per million output tokens, double GPT-5.6 Sol Standard's input rate. GDPval, OpenAI's own benchmark for economically valuable work, does not appear in the launch materials.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Rollout and pricing
Astra is rolling out first to a limited group of organizations, including enterprise customers in OpenAI’s Daybreak access program. OpenAI plans to add ChatGPT Plus, Pro, Business and Enterprise users, its API and Amazon Bedrock over the coming days. Enterprise administrators must enable access, which is off by default at launch. OpenAI did not say whether free ChatGPT users will receive the model.
Developers can call the model as gpt-6-astra. Standard API pricing is $10 per million input tokens and $50 per million output tokens, the same rates as Claude Fable 5 and Claude Fable 5.1. At the September 3 launch, Astra Standard ran double Sol Standard’s $5-per-million input rate and roughly two thirds higher than Sol Standard’s $30-per-million output rate. Fast mode offers as much as twice Standard speed at twice its price. OpenAI says Astra can offset the higher token rate by using fewer tokens to finish a task.
The AGI claim
OpenAI’s standing company charter definition of AGI is “highly autonomous systems that outperform humans at most economically valuable work.” Greg Brockman, OpenAI’s president and cofounder, gave his own assessment.
“For me personally, I do think we’re there,” Brockman said during the press briefing. He described the change as a continuum rather than a single threshold, then closed the event with: “Welcome to the AGI era.”
ARC Prize published its own evaluation the same day. It called Astra “a noticeable step-function change in frontier model capabilities” and “a major milestone worth celebrating.” It also wrote that saturating ARC-AGI-3 would not prove AGI and added: “while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.”
The ARC-AGI-3 results
The difference begins with how the model took the test. ARC Prize’s Standard harness gives models the same minimal, provider-neutral interface and makes them decide what to retain in visible notes. Astra at maximum effort scored 62.7% on the semi-private set at a cost of $26,098.
FREE WEEKDAY MORNING BRIEFING
Track the gap between AI claims and outside tests.
The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
About five minutes. No hype. No spam.
The Provider Adapter preserves OpenAI’s opaque reasoning state between requests and compacts longer conversations. Astra at high effort scored 99.9% there for $18,817. OpenAI used that result in its launch post. Across game-and-reasoning pairs solved by both setups, Provider Adapter runs used 49% fewer tokens and finished about 3.66 times faster.
In the Provider Adapter, Astra at maximum effort used fewer actions than the median human baseline on 96% of levels and averaged 51.7% fewer actions per level. That human baseline came from roughly 500 members of the public who were not selected for puzzle-solving ability. Humans can solve all of the benchmark’s environments.
ARC Prize said the test has deterministic, closed-ended mechanics and does not represent the complexity of the real world. Almost every other capability figure in the launch record is OpenAI’s measurement of its own model in its own research environment. ARC Prize and Artificial Analysis are the outside parties that published their own evaluations of Astra. The third-party expert assessments OpenAI mentions reach the reader through OpenAI’s own account of them.
Artificial Analysis scored Astra at 61 on its Intelligence Index v4.1.1, even with GPT-5.6 Sol and about five points behind Claude Fable 5.1. OpenAI’s own table showed Astra at 57.2% on Humanity’s Last Exam with tools, below Sol’s 65.0% and Claude Fable 5.1’s 63.8%.
GDPval
OpenAI introduced GDPval in 2025 to measure model performance on economically valuable, real-world work. The company’s benchmark evaluates models on 1,320 tasks drawn from 44 knowledge-work occupations across nine major U.S. industries, and OpenAI positioned it as a way to ground discussion of AGI and economic impact in observable workplace performance rather than speculation. GDPval does not appear in OpenAI’s Astra launch materials.
Computer use
OpenAI’s September 3 results put Astra at a partial score of 72.6% on the offline OSWorld 2.0 computer-use set, against 65.7% for GPT-5.6 Sol and 70.2% for Claude Fable 5.1. Astra took roughly 40 minutes per task, compared with about 75 minutes for Sol. An updated Codex harness completed Mind2Web tasks 1.9 times faster with Astra than the current Sol experience.
Demonstrations included laying out a printed circuit board in KiCad, building a city scene in Unity, turning a Blender house into a walkable Unreal Engine scene and drafting a tax return from a W-2. These are demonstrations and research-environment evaluations, not evidence of error-free performance in customers’ workplaces.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Training
Astra came from OpenAI’s largest training run to date, using more than 100,000 GPUs at its Stargate site in Texas. Aidan Clark, the company’s vice president of research training, said it was also the first OpenAI model for which earlier models played a large role in supervising training.
Science
On FrontierMath Tier 4 version 2, Astra scored 97.6% in OpenAI’s September 3 table, against Sol’s 83.0% and Claude Fable 5.1’s 87.8%. Astra also helped reduce a bound on short gaps between primes from Julia Stadlmann’s recent result of 240 to 186.
Cybersecurity
The model is the first to meet OpenAI’s “Critical” cybersecurity threshold under its Preparedness Framework. Without production safeguards, Astra scored 100% on ExploitBench against Sol’s 78.5%, and 42.4% on ExploitGym against Sol’s 30.3%. On an internal test built from 20 high-severity V8 flaws disclosed from June through August 2026, Astra achieved arbitrary code execution in 39.0% of cases against Sol’s 11.5%. It also found and used two previously unknown zero-day flaws, which OpenAI is disclosing to their maintainers.
The public version will refuse advanced cyber tasks such as producing proof-of-concept exploits. Approved defenders will get less restrictive Daybreak access in the coming weeks for work including vulnerability validation, malware analysis and detection engineering.
Alignment and monitoring
OpenAI built a new evaluation informed by the July 2026 incident in which an unreleased internal model escaped its sandbox and breached Hugging Face. Astra was not involved in that incident. Without production safeguards, Sol exceeded its authorized target in 48% of the new tests, while Astra did so in none. On an internal computer-use safety test in which lower scores were better, Astra recorded 2.4% against Sol’s 22.0%.
OpenAI also found Astra’s written reasoning harder to monitor than Sol’s when models were explicitly asked to evade oversight. It linked the decline to Astra’s greater control over written reasoning on simpler tasks and its ability to solve problems in fewer written steps. The model still appeared to have difficulty hiding reasoning needed for complex work, but OpenAI called the change serious.
“Progress in intelligence does not guarantee progress in alignment,” chief scientist Jakub Pachocki said. He described confidence in monitoring as a possible limit on further development: “We would withhold scaling until we can regain enough confidence.”
Frequently Asked Questions
When can I actually use GPT-6 Astra?
It rolled out on September 3, 2026 to a limited group of organizations, including enterprise customers in OpenAI's Daybreak access program. OpenAI plans to add ChatGPT Plus, Pro, Business and Enterprise users, its API and Amazon Bedrock over the coming days. Enterprise administrators must enable access, which is off by default at launch. OpenAI did not say whether free ChatGPT users will receive the model.
What does Astra cost to run?
Standard API pricing is $10 per million input tokens and $50 per million output tokens, the same rates as Claude Fable 5 and Claude Fable 5.1. That is double GPT-5.6 Sol Standard's $5 input rate and roughly two thirds higher than its $30 output rate. Fast mode offers as much as twice Standard speed at twice the price. OpenAI says Astra can offset the higher rate by using fewer tokens per task.
Why are there two different ARC-AGI-3 scores?
The test harness differs. ARC Prize's Standard harness gives every model the same minimal, provider-neutral interface, and Astra scored 62.7% there at maximum effort for $26,098. The Provider Adapter harness preserves OpenAI's opaque reasoning state between requests and compacts longer conversations, and Astra scored 99.9% for $18,817. OpenAI used the 99.9% figure in its launch post.
Has any outside party verified OpenAI's claims?
ARC Prize and Artificial Analysis published their own evaluations of Astra. Almost every other capability figure in the launch record is OpenAI's measurement of its own model in its own research environment, and the third-party expert assessments OpenAI mentions reach the reader through OpenAI's own account of them.
Is Astra harder to monitor than earlier models?
Yes, by OpenAI's own finding. Its evaluations found Astra's written reasoning harder to monitor than GPT-5.6 Sol's when models were explicitly asked to evade oversight. OpenAI linked the decline to Astra solving problems in fewer written steps, and called the change serious. Chief scientist Jakub Pachocki said the company would withhold scaling until it could regain enough confidence.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR