OpenAI Says Its Upcoming Astra Model Is The First To Cross The Critical Cyber Threshold

OpenAI told the safety community it can no longer rule out that its upcoming Astra model has crossed the Critical cybersecurity threshold in its Preparedness Framework after Astra chained two zero-days without human guidance.

OpenAI Says Its Upcoming Astra Model Is The First To Cross The Critical Cyber Threshold

OpenAI told the safety and security community on September 1, 2026 that it can no longer rule out its upcoming Astra model having crossed the Critical cybersecurity threshold in its Preparedness Framework. The disclosure, published in a blog post titled “Responding to the next frontier of critical cyber capabilities,” marks the first time any OpenAI model has forced the company to escalate its own safeguards to the top of the ladder.

What The Critical Threshold Actually Means

Under the framework, a model reaches Critical if it can identify and develop working zero-day exploits across hardened real-world systems without human intervention, or if it can devise and execute end-to-end novel cyberattacks against hardened targets given only a high-level goal. During internal testing, Astra reportedly discovered and chained together two previously unknown vulnerabilities, which OpenAI says it is disclosing to the affected maintainers. Previous models, including GPT-5.6-Sol, only reached the High tier.

OpenAI stressed that Astra had no role in the recent Hugging Face incident, addressed separately in its post-mortem.

OpenAI Path to Astra cover

New Controls Around A More Dangerous Model

OpenAI paused internal Astra activities that fell short of the reinforced controls and rolled out isolated testing environments, restricted network and tool access, enhanced model-weight encryption, sandboxed execution and universal monitoring across every agentic run. Chain-of-thought monitors now trigger a security response whenever they flag high-risk behaviour, and the company says it will hand external testing over to relevant government agencies and select AI safety organisations before wider deployment.

Race Against Adversarial AI Just Tightened

The disclosure lands the same week CrowdStrike and NVIDIA launched SafeMind, whose whole design brief is to pit offence against defence in a continuous loop precisely because attackers are gaining access to frontier tooling. Astra will still ship, OpenAI said, but with tighter throttles on its most advanced cybersecurity capabilities and a Defender program that steers those abilities into vulnerability discovery for the good side of the ledger.

Reporting based on coverage from OpenAI, Axios and PYMNTS.

Category: Machine Learning

Tags: Cybersecurity AI Foundation Models OpenAI agentic AI AI Red Teaming AI safety

Related Articles