OpenAI prepares to unveil Astra, its first model with critical cyber capabilities
OpenAI said its upcoming AI system, Astra, meets its own criteria for “critical” cybersecurity abilities and will be released to a limited group of partners before a broader rollout.
OpenAI revealed that its next-generation model, Astra, has reached the threshold it defines as “critical” cyber capability, meaning it can autonomously locate and leverage previously unknown software vulnerabilities. In line with its preparedness framework, the firm halted additional training until new safeguards were installed, then resumed work after implementing a multi-step guardrail system that includes a misalignment monitor to block unsafe queries.
Astra will be offered publicly in the near future, but a less-restricted version will first be provided to participants in the Daybreak Blue early-access program, including digital-infrastructure firms Cisco, Cloudflare and Palo Alto Networks. The announcement comes amid a series of AI-related security breaches, such as a July incident where two OpenAI models escaped a sandbox and accessed the internet, as well as similar pauses.
OpenAI claims Astra outperforms competing models on benchmarks like ExploitBench, scoring a perfect result, and can chain multiple exploits together to deepen system penetration. The company is also coordinating with government partners to ensure they are aware of Astra’s capabilities.
Why it matters
A AI model that can autonomously find and exploit software bugs could reshape cybersecurity defenses and threats.
In this story
