GPT-Red improves model robustness through automated red-teaming

GPT-Red red-teaming model helps OpenAI find AI vulnerabilities before launch Boost prompt-injection defense with scalable testing and stronger safety checks

OpenAI says it has developed GPTRed, an internal automated redteaming model designed to find vulnerabilities in its AI systems and help improve safety before deployment. The company says the approach is meant to scale redteaming beyond what human testers can do alone, especially as models become more capable and interact with thirdparty data through tools like browsers, apps, and files. According to the post, GPTRed is trained through selfplay and is used to generate adversarial examples for newer production models, including GPT5.6 Sol. OpenAI says this has made its latest model more resistant to prompt injection attacks, with fewer failures on its hardest direct prompt injection benchmark than earlier production models. The company also describes tests in which GPTRed was able to challenge novel environments and realworldstyle agent systems, including an AIpowered vending machine and a Codex CLI agent. OpenAI says the goal is not to deploy GPTRed itself, but to use it to strengthen future models while continuing human redteaming, thirdparty testing, layered safeguards, and monitoring.