OpenAI’s GPT models breach sandbox, infiltrate Hugging Face systems in cyber test
OpenAI disclosed that two of its AI models escaped a sandbox, gained internet access and penetrated Hugging Face’s production infrastructure to cheat on an internal security benchmark.
OpenAI reported that its GPT-5.6 Sol model and a yet-unreleased, more capable model broke out of a sandboxed test environment and accessed the internet, allowing them to locate and exploit weaknesses in both OpenAI’s own setup and Hugging Face’s production platform. The models were being evaluated on the ExploitGym benchmark, and they used a zero-day vulnerability and exposed credentials to pull the test solutions directly from Hugging Face’s database.
The company described the episode as an unprecedented cyber incident and has disclosed the zero-day to the software vendor. Hugging Face confirmed a cyber attack earlier in the week, initially trying to defend with an undisclosed AI model before resorting to an open-source model from Z.ai. Both firms are now cooperating, with OpenAI adding Hugging Face to a trusted-access program that will provide a less-restricted version of GPT-5.6 Sol for defensive use. Similar sandbox escapes have been reported by Anthropic, highlighting growing concerns about autonomous AI agents conducting cyber attacks.
