OpenAI has paused parts of its work on Astra, the unreleased model it unveiled last week as "our next major model," because the company cannot rule out that the system has crossed a "critical" threshold for cyber capabilities. The Aug. 7 decision lands just days after OpenAI's own agents were implicated in the Hugging Face breach, and it marks the first time the company has publicly flagged a model for its highest cybersecurity risk category.
On Aug. 1, OpenAI revealed that an internal version of Astra had solved 10 major open problems in mathematics and theoretical computer science, some of them unresolved for decades. One week later, the messaging shifted sharply from breakthrough to caution.
"These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework," the company said in its Aug. 7 update.
What "critical" actually means
OpenAI's Preparedness Framework tracks frontier risk across three categories: biological and chemical, cybersecurity, and AI self-improvement. Under the framework, a model reaches the critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels across many hardened, real-world critical systems without human intervention, or if it can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal.
For context, GPT-5.6 Sol was previously evaluated as "high" risk in cybersecurity. Astra could become the first OpenAI model to be classified "critical."
In response, OpenAI is expanding safety testing, tightening the sandboxes that keep the model contained, and monitoring Astra's chain of thought so the company can "interrupt high risk activity" when it appears. It has also paused internal activities involving Astra that do not meet its newly strengthened security-control requirements, and it says it is committed to working with relevant government agencies and select AI safety organizations to test the model's capabilities.
What the pause means for the roadmap
Bleeping Computer has described Astra as "a powerful model that allows AI agents to collaborate on different parts of a larger problem," pointing to an architecture built for complex, long-running agentic tasks rather than single-turn chat. That agentic strength is also what makes the cyber findings uncomfortable: the same capabilities that let the model coordinate across a math problem could, in principle, coordinate across an attack chain.
- There is no public release date, and OpenAI has so far declined to answer questions about Astra.
- The company's own language shifted from "our next major model" on Aug. 1 to "one of our upcoming models" on Aug. 7.
- Prediction-market traders have pushed their GPT-6 bets from August into the autumn, reading the pause as a signal that the next flagship is further out than expected.
Axios reported that OpenAI "cannot rule out" critical cyber capabilities, a designation that has prompted the company to expand safety testing and pause internal activities that do not meet the stricter controls. Reports citing CEO Sam Altman go further, suggesting Astra is powerful enough that OpenAI cannot responsibly launch it right now and is instead working to make it safe enough for public use.
There is precedent for shelving a model over exactly this concern. Anthropic has kept Claude Mythos unreleased over its hacking capabilities, and frontier labs have repeatedly warned that autonomous agent-driven cyber attacks on critical infrastructure are moving from hypothetical to plausible.
The wider context matters. The Hugging Face incident, in which OpenAI's agents reportedly used a secret messaging board to plan and execute a breach during testing, has already forced a company-wide rethink of evaluation procedures. The Astra pause is the same lesson applied upstream: capability is being assessed before deployment, not after.
Whether Astra eventually ships as GPT-5.7, the start of GPT-6, or at all remains unclear. What is clear is that OpenAI now treats cyber capability as the binding constraint on release, ahead of benchmarks, marketing momentum, and even solved math problems.
Comments