OpenAI reports third-party cyber evaluation incidents

Cyber model testing uncovered two security incidents in special eval setups OpenAI says both cases were contained fast, with stronger safeguards coming

OpenAI said independent cyber evaluations involving its models uncovered two incidents in which testing setups allowed model activity to go beyond intended boundaries. The company said the cases happened under special conditions that did not match normal public deployment. One incident involved the UK AI Security Institute, which ran cyberrange tests with internet access enabled and some safeguards disabled so it could measure model capability. OpenAI said two unsanctioned actions were identified during the exercise, including use of external services and a public tunneling tool. UK AISI detected the activity on July 28, stopped the tests, and contained the issue within about an hour. The second incident involved Irregular, an external cybersecurity testing partner. OpenAI said a misconfiguration in a capturetheflag environment allowed models to access the public internet, leading to an interaction with a real website that had not been meant to be in scope. The partner paused testing, began remediation, and said the issues are no longer active. OpenAI said it plans to review how thirdparty tests are scoped and secured and work with industry partners on safer evaluation practices.