Anthropic details Fable 5 cyber safeguards and jailbreak framework

Claude Fable 5 cyber safeguards: see how Anthropic blocks risky requests and why its jailbreak severity framework could shape AI security standards

Anthropic has released additional information about the cybersecurity safeguards built into Claude Fable 5, along with an early draft of a proposed framework for rating AI jailbreak severity. The company said the model uses safety classifiers to detect and block dangerous or potentially dangerous cyber requests, while still allowing some benign and defensive uses. Anthropic organized cyberrelated behavior into four categories: prohibited use, highrisk dual use, lowrisk dual use, and benign use. It said the classifiers are only one layer of protection, alongside access controls, safety training, and offline monitoring. Anthropic also outlined a Cyber Jailbreak Severity scale meant to help describe how serious a jailbreak is based on factors such as the capability it unlocks, how broadly it applies, how easy it is to weaponize, and how discoverable it is. The company said the framework is still a draft and is intended to encourage discussion with researchers, industry, and government on how to assess these risks consistently.