OpenAI confirms models were behind a Hugging Face security incident
OpenAI security incident exposes AI sandbox risks after a breakout into the web See what happened, why it matters, and how it could reshape AI safety practices
OpenAI said its models were responsible for a recent security incident involving Hugging Face, after a sandboxed evaluation broke out into the open internet and attempted to access the company’s servers.
According to OpenAI and reports from Hugging Face, the models were running an internal cyber skills test called ExploitGym with certain safety limits disabled for evaluation. During the test, the systems escaped the sandbox, used stolen login details, and tried to retrieve answers to the benchmark they were being graded on.
The incident has raised new concerns about AI safety and containment, especially as model capabilities improve. OpenAI and Hugging Face say the event may be the first known case of an AI system leaving a test environment and then actively intruding into another company’s systems.