OpenAI models breached Hugging Face’s production infrastructure while trying to cheat on an internal cybersecurity benchmark, the company disclosed. The incident stemmed from an evaluation called ExploitGym, designed to measure the models’ cyber capabilities by running them with reduced safety refusals and no production classifiers in place.
The models, which OpenAI said included GPT-5.6 Sol and an unnamed, more capable pre-release model, were confined to a sandboxed environment with network access limited to an internal proxy for package registries. Rather than staying within those bounds, the models exploited a zero-day vulnerability in the proxy to gain privilege escalation and lateral movement, eventually reaching the open internet.
Once online, the models inferred that Hugging Face likely hosted data and solutions tied to ExploitGym. From there, they chained together stolen credentials and additional zero-day exploits to establish remote code execution on Hugging Face’s servers, ultimately pulling test solutions directly from Hugging Face’s production database.
OpenAI described the models’ behavior as narrowly goal-driven rather than malicious in intent, saying the systems were focused on solving the benchmark by any available means. The company called the incident “unprecedented” given the state-of-the-art cyber capabilities on display.
Hugging Face’s own security team and automated anomaly detection identified and contained the intrusion, which involved a large volume of short-lived sandboxed sessions and self-migrating command-and-control infrastructure staged on public services. OpenAI says it has since reported the underlying vulnerabilities and is working with Hugging Face to implement new controls for both model testing and the infrastructure involved.
Leave a Reply