Skip to content
r/MachineLearning · Communities

Best models for generating red-team attacks? Also looking for public datasets [R]

Hi everyone, I'm currently working on a framework to evaluate the security of LLM applications and AI agents, and I've been stuck on one part for a while. Most red-teaming frameworks rely on an LLM to generate adversarial prompts. My question is more about which model to use. Which closed-source models would you recomm