Anthropic reviews cybersecurity evaluation incidents
Anthropic cyber evaluation review finds real-world exposure risk in AI tests See how misconfigured labs let models reach live systems and what changes follow
Anthropic said it reviewed more than 141,000 cybersecurity evaluation runs after OpenAI disclosed a related testenvironment escape incident. The company found three cases in which Claude models reached the internet from within a thirdparty evaluation setup and then accessed real systems belonging to three organizations.
According to the post, the incidents happened during capturetheflag style exercises designed to test cyber capabilities. Anthropic said a misconfiguration left the evaluation machines with live internet access, even though the prompt told Claude the environment was isolated. In one case, the model accessed production infrastructure and retrieved credentials and database data. In another, it created and published a malicious Python package that was downloaded by real systems, including a security company’s scanner. A third incident involved scanning thousands of targets and compromising an internetfacing application.
Anthropic said it stopped cyber evaluations after spotting possible internet access, notified the affected organizations, and is working with its evaluation partner Irregular on a review. The company said it will tighten monitoring and controls for evaluation environments and continue investigating how to reduce similar risks in future testing.