Skip to content
HF Daily Papers · Papers

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust mo