OpenAI Pauses Astra AI Model Over Critical Cyber Capabilities
OpenAI announced it is pausing certain internal activities related to its upcoming artificial intelligence model, Astra, after internal evaluations revealed the model had achieved significant advancements in agentic coding and cybersecurity capabilities. The decision, disclosed on August 10, 2026, stems from preliminary assessments indicating that Astra's performance in cyber-related tasks is "strong enough" that OpenAI cannot rule out the model reaching "Critical" capability status under its Preparedness Framework.
In response, OpenAI is rolling out enhanced security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. The company stated it is "pausing internal activities involving Astra that do not yet meet these strengthened security control requirements." OpenAI has also implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. These monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high-risk activity in real time.
Under OpenAI's Preparedness Framework, "Critical" cyber capability is defined as a tool-augmented model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, OR that can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. OpenAI emphasized that preliminary evaluations could not eliminate the possibility that Astra meets this threshold. The company separately clarified that Astra was not involved in last month's security incident targeting Hugging Face. In a recent academic paper, OpenAI also noted that Astra solved 10 open problems in mathematics and theoretical computer science for approximately $2,000 at Sol API rates.
OpenAI stated it will collaborate with relevant government agencies and select AI safety organizations to safely test Astra's capabilities, while also sharing recommended security controls with third-party testing partners to run higher-risk evaluations. As agentic AI models grow more capable, defenders should harden their own environments: run a port scanner to identify exposed services, verify certificate hygiene with an SSL/TLS checker, and validate account credentials using a password checker. OpenAI affirmed its position on transparency, stating, "We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do," and pledged continued work alongside governments, safety institutes, and civil society.