News

OpenAI Restricts Astra Work After Cybersecurity Warning

OpenAI has restricted some internal work involving an unreleased artificial intelligence model after preliminary evaluations indicated it might possess cybersecurity capabilities at the highest level defined by the company’s safety framework.

The company said on Aug. 7 that recent testing of the model, known as Astra, showed advances in agentic coding and cybersecurity. The results led OpenAI to conclude that it “cannot rule out critical cyber capabilities” under its Preparedness Framework.

The statement does not amount to a formal determination that Astra has reached the Critical threshold. OpenAI said its findings were preliminary and that benchmarking and expert assessment were continuing.

The company also has not stopped all development of Astra. It said it is pausing internal activities involving the model when those activities do not meet strengthened security requirements. Work performed under the new controls can continue.

The distinction is important because the announcement describes uncertainty about an emerging capability, not evidence that Astra has conducted a successful attack against a real target.

A threshold designed for autonomous attacks
OpenAI’s Preparedness Framework covers frontier capabilities that the company believes could create risks of severe harm. Its tracked categories include cybersecurity, biological and chemical capabilities, and AI self-improvement.

Under the framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits across many hardened, real-world critical systems without human assistance. A model could also qualify if it can devise and execute a novel, end-to-end attack against a hardened target after receiving only a high-level objective.

That standard is higher than helping an experienced operator write malicious code or solve controlled security exercises. It is intended to capture capabilities that could create a new route to severe harm, rather than merely making existing attacks somewhat easier.

The framework defines High capabilities as those that significantly increase existing risk. Critical capabilities represent a potentially new threat without a ready precedent.

OpenAI’s previously evaluated models, including GPT-5.6 Sol, were classified at High for cybersecurity rather than Critical. The company’s August update for the GPT-5.6 family said Sol and Luna were being treated as High in both cybersecurity and biological and chemical capabilities.

Development controls, not only deployment safeguards
OpenAI’s framework requires safeguards before a model rated High can be deployed. If a model under development reaches the Critical threshold, safeguards must also cover the development process, regardless of whether the company intends to release it.

In response to the Astra results, OpenAI said it is implementing isolated testing environments, restricted network access, and stronger protection for model weights. It also described additional monitoring and sandboxed execution.

The company said all agentic applications of Astra, including training and evaluation, are subject to monitoring for risky behavior and possible misalignment. Those systems can trigger a security review and interrupt activity judged to present a high risk.

OpenAI said it would work with government agencies and selected AI safety organizations to test the model. It also plans to provide security recommendations to outside partners conducting higher-risk evaluations.

The measures reflect a problem created by agentic models. A conventional chatbot generally responds to a prompt with information. An agent can be connected to software tools, command-line environments, or networks, allowing it to take actions over multiple steps. That makes the security of the surrounding system relevant alongside the model's behavior.

Evidence remains limited
OpenAI has not published the evaluations that triggered its response. The company has not disclosed Astra’s architecture, parameter count, or intended product role. It has also not provided a release date.

No public system card identifies the tasks Astra completed or explains how close its results were to OpenAI’s Critical threshold. The company referred to internal evaluations and expert assessments but did not name the outside experts involved.

That leaves independent researchers unable to assess whether Astra demonstrated a broad capability or performed strongly on a smaller number of tests.

OpenAI also said Astra was not involved in exploiting Hugging Face, referring to a recent security incident. The company did not indicate that Astra had been used in any other real-world intrusion.

The claim that this is the first time an AI laboratory has flagged one of its models at a Critical cyber threshold is not established by OpenAI’s announcement. It may be the first such public warning issued under OpenAI’s framework, but other laboratories use different assessment systems and may not disclose internal results.

The dual-use problem
More capable cybersecurity models can help defenders identify vulnerabilities and produce fixes. The same abilities could also assist attackers in finding weaknesses or automating parts of an intrusion.

OpenAI has sought to manage that tension through its Trusted Access for Cyber program, which gives vetted defenders access to more permissive cybersecurity tools. The company says the program uses identity verification and tiered access to support legitimate work while restricting requests that could facilitate real-world harm.

Astra could test whether that approach remains workable as models become more autonomous. Access controls can limit who uses a deployed service, but they do not by themselves address risks during training, internal testing, or the handling of model weights.

OpenAI’s Preparedness Framework places responsibility for reviewing such risks with its Safety Advisory Group, an internal body that recommends safeguards and deployment decisions. Company leadership makes final decisions, while the board’s Safety and Security Committee has an oversight role.

For now, Astra remains an unreleased model undergoing additional evaluation. OpenAI has not said whether the new controls will delay a planned launch or whether Astra will eventually receive a formal Critical designation.

The immediate significance is narrower. A safety threshold written to govern future models is now influencing work on an actual system under development. Whether Astra crosses that threshold remains unresolved.

About the Author

John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS.  He can be reached at [email protected].

Featured