OpenAI says Astra meets critical cybersecurity threshold

OpenAI Astra meets critical cybersecurity threshold with stronger safeguards Early access is limited while OpenAI expands defenses and shares safety results

OpenAI said on September 1, 2026, that its model Astra now meets a critical cybersecurity capability threshold under the company’s Preparedness Framework. The company said the model can identify unknown security flaws and develop exploit paths when given the right tools and access, which prompted stronger safeguards before release. The company said it delayed parts of Astra’s development while it strengthened protections against cyber misuse and unauthorized actions. OpenAI added that it has increased refusal training, monitoring, and other controls to reduce the risk of severe harm, and said its safeguards are designed to block both malicious use and any misaligned actions by the model itself. Astra’s most advanced cybersecurity features will not be broadly available at launch. OpenAI said initial access will go to a small group of testers, with broader defensive access planned later through Daybreak Blue. The company said it will publish more detailed safety and evaluation results in the model’s system card at launch.