OpenAI Details for GPT-Red: An Automated Internal Model of the Red Team Beats Human Red Teams 84% to 13% in Rapid Injection
This week, OpenAI published details of the GPT-Redthe only internal model that defaults to the red group. Its job is to attack OpenAI models and detect rapid injection vulnerabilities. OpenAI offers two reasons. Bringing people together in red takes time and is not fair. Commonly used durability tests are now available in the latest models. … Read more