AI-Assisted Red-Teaming
As AI system technology advances, it increasingly exhibits emergent behaviors—capabilities that are not explicitly programmed or apparent in earlier models. For example, GPT-3 has demonstrated the ability to write code snippets, despite not being trained at scale on code like its successor, Codex. Security operators cannot fully guard against potentially harmful emergent behaviors, since these behaviors are, by definition, unprecedented and unexpected.
Red-teaming is one approach to testing AI models for such harmful behaviors, including the leakage of sensitive data and the generation of toxic, biased, or factually inaccurate content. The goal of red-teaming is to probe for latent capabilities and potential misuse pathways, rather than focusing solely on the accuracy or toxicity of LLM outputs.
Commercial Off-the-Shelf Solutions and Simple Recommendations
- Microsoft AI Red Team1
- For known vulnerabilities, see the Open Vulnerability Assessment Scanner (OPENVAS)2 and Nessus vulnerability scanner.3
- Some open-source tools are the Python Risk Identification Tool (PyRIT),4 DeepTeam,5 Garak,6 and Woodpecker.7
See also
- AdverTorch8
- AI Fairness 3609
- ART10
- BrokenHill11
- BurpGPT12
- CleverHans13
- Counterfit14
- Crucible by Dreadnode15
- Foolbox16
- Galah17
- Gepetto18
- Ghidra19
- GPT-WPRE20
- Granica21
- Guardrails-AI22
- HiddenLayer23
- IATelligence24
- Inspect25
- Jailbreak-evaluation26
- LLMFuzzer27
- LM Evaluation Harness28
- Meerkat29
- Mend.io30
- Mindgard31
- Plexiglass32
- PowerPwn33
- Purple Llama34
- SecML35
- TextAttack36
- ThreatModeler AI37
- Vigil38
How Does This Relate to the Rest of the Guide or Other Threats That the User Cares About?
- Knowledge hazard: A.PPT-1, D.RDTP-1, D.RDTP-2
- Red-teaming: MT.ART-1
- OWASP:
- API1:2023—Broken Object Level Authorization
- LLM01: Prompt Injection
- ML01: Input Manipulation Attack
Other Sources of Information About This Topic
- CISO Guide (Breachlock)39
- “31 Best Tools for Red Teaming” (Mindgard)40
- “Best AI Red Teaming Tools” (Mend.io)41
- “Emergent Behavior in Large Language Models” (ThirdEye)42
- “What Is Red Teaming for Generative AI?” (IBM)43
- “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence” (White House)44
- “AI Assisted Red Teaming—an Efficient Approach” (Medium)45
- “What Does AI Red-Teaming Actually Mean?” (CSET)46
- “A Guide to AI Red Teaming” (HiddenLayer)47
- LLM and Generative AI Security Solutions Landscape—Q1,2025 (OWASP)48
- “Are Emergent Abilities in Large Language Models Just In-Context Learning?” (arXiv)49
- “Emergent Abilities in Large Language Models” (arXiv)50
- “Context-Aware LLM-Based Safe Control Against Latent Risks” (arXiv)51
- “Connecting the Dots” (Neural Information Processing Systems Foundation, Inc.)52
- “AART” (arXiv)53
- “Curiosity-Driven Red-Teaming for Large Language Models” (arXiv)54
Notes
- Microsoft, “Microsoft AI Red Team.” Return to content ⤴
- OPENVAS, homepage. Return to content ⤴
- Tenable, “Tenable Nessus®.” Return to content ⤴
- Microsoft Azure, “PyRIT.” Return to content ⤴
- DeepTeam, homepage. Return to content ⤴
- NVIDIA, “Garak.” Return to content ⤴
- OperantAI, “Woodpecker.” Return to content ⤴
- BorealisAI, “AdverTorch.” Return to content ⤴
- AI Fairness 360, homepage. Return to content ⤴
- Guanlin Lee, “ART.” Return to content ⤴
- Bishop Fox, “BrokenHill.” Return to content ⤴
- BurpGPT, homepage. Return to content ⤴
- CleverHans Lab, “Cleverhans.” Return to content ⤴
- Microsoft Azure, “Counterfit.” Return to content ⤴
- Dreadnode, homepage. Return to content ⤴
- Bethge Lab, “Foolbox.” Return to content ⤴
- Ka, “Galah.” Return to content ⤴
- Kwiatkowski, “Gepetto.” Return to content ⤴
- Tenable, “Ghidra_tools.” Return to content ⤴
- Dolan-Gavitt, “GPT-WPRE.” Return to content ⤴
- Granica, homepage. Return to content ⤴
- Guardrails AI, homepage. Return to content ⤴
- Smith, “A Guide to AI Red Teaming.” Return to content ⤴
- Roccia, “IATelligence.” Return to content ⤴
- AI Security Institute, “Inspect AI.” Return to content ⤴
- Controllability, “Jailbreak-Evaluation.” Return to content ⤴
- Brin, “LLMFuzzer.” Return to content ⤴
- EleutherAI, “LM-Evaluation-Harness.” Return to content ⤴
- HazyResearch, “Meerkat.” Return to content ⤴
- Mend.io, homepage. Return to content ⤴
- Mindgard, “Offensive Security Testing for Your AI”; Glynn, “What Is AI Red Teaming?” Return to content ⤴
- SafeLlama, “Plexiglass.” Return to content ⤴
- Bargury, “Power-pwn.” Return to content ⤴
- Meta, “Introducing Purple Llama for Safe and Responsible AI Development.” Return to content ⤴
- Pattern Recognition and Applications Lab and Pluribus One, “SecML.” Return to content ⤴
- TextAttack, homepage. Return to content ⤴
- ThreatModeler, homepage. Return to content ⤴
- Swanda, “Vigil-LLM.” Return to content ⤴
- BreachLock, “CISO Guide.” Return to content ⤴
- Glynn, “31 Best Tools for Red Teaming.” Return to content ⤴
- Tayouri and Haas, “Best AI Red Teaming Tools.” Return to content ⤴
- ThirdEye Data, “All About Emergent Behavior in Large Language Models.” Return to content ⤴
- Martineau, “What Is Red Teaming for Generative AI?” Return to content ⤴
- Executive Order 14110, “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence.” Return to content ⤴
- Roshan, “AI Assisted Red Teaming—an Efficient Approach.” Return to content ⤴
- Ji, “What Does AI Red-Teaming Actually Mean?” Return to content ⤴
- Smith, “A Guide to AI Red Teaming.” Return to content ⤴
- OWASP, LLM and Generative AI Security Solutions Landscape—Q1,2025. Return to content ⤴
- Lu et al., “Are Emergent Abilities in Large Language Models Just In-Context Learning?” Return to content ⤴
- Berti, Giorgi, and Kasneci, “Emergent Abilities in Large Language Models.” Return to content ⤴
- Deng et al., “Context-Aware LLM-Based Safe Control Against Latent Risks.” Return to content ⤴
- Treutlein et al., “Connecting the Dots.” Return to content ⤴
- Radharapu et al., “AART.” Return to content ⤴
- Hong et al., “Curiosity-Driven Red-Teaming for Large Language Models.” Return to content ⤴