AI-Assisted Red-Teaming

As AI system technology advances, it increasingly exhibits emergent behaviors—capabilities that are not explicitly programmed or apparent in earlier models. For example, GPT-3 has demonstrated the ability to write code snippets, despite not being trained at scale on code like its successor, Codex. Security operators cannot fully guard against potentially harmful emergent behaviors, since these behaviors are, by definition, unprecedented and unexpected.

Red-teaming is one approach to testing AI models for such harmful behaviors, including the leakage of sensitive data and the generation of toxic, biased, or factually inaccurate content. The goal of red-teaming is to probe for latent capabilities and potential misuse pathways, rather than focusing solely on the accuracy or toxicity of LLM outputs.

Commercial Off-the-Shelf Solutions and Simple Recommendations

  • Microsoft AI Red Team⁠1
  • For known vulnerabilities, see the Open Vulnerability Assessment Scanner (OPENVAS)⁠2 and Nessus vulnerability scanner.⁠3
  • Some open-source tools are the Python Risk Identification Tool (PyRIT),⁠4 DeepTeam,⁠5 Garak,⁠6 and Woodpecker.⁠7

See also

How Does This Relate to the Rest of the Guide or Other Threats That the User Cares About?

  • Knowledge hazard: A.PPT-1, D.RDTP-1, D.RDTP-2
  • Red-teaming: MT.ART-1
  • OWASP:
    • API1:2023—Broken Object Level Authorization
    • LLM01: Prompt Injection
    • ML01: Input Manipulation Attack

Other Sources of Information About This Topic

  • CISO Guide (Breachlock)⁠39
  • “31 Best Tools for Red Teaming” (Mindgard)⁠40
  • “Best AI Red Teaming Tools” (Mend.io)⁠41
  • “Emergent Behavior in Large Language Models” (ThirdEye)⁠42
  • “What Is Red Teaming for Generative AI?” (IBM)⁠43
  • “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence” (White House)⁠44
  • “AI Assisted Red Teaming—an Efficient Approach” (Medium)⁠45
  • “What Does AI Red-Teaming Actually Mean?” (CSET)⁠46
  • “A Guide to AI Red Teaming” (HiddenLayer)⁠47
  • LLM and Generative AI Security Solutions Landscape—Q1,2025 (OWASP)⁠48
  • “Are Emergent Abilities in Large Language Models Just In-Context Learning?” (arXiv)⁠49
  • “Emergent Abilities in Large Language Models” (arXiv)⁠50
  • “Context-Aware LLM-Based Safe Control Against Latent Risks” (arXiv)⁠51
  • “Connecting the Dots” (Neural Information Processing Systems Foundation, Inc.)⁠52
  • “AART” (arXiv)⁠53
  • “Curiosity-Driven Red-Teaming for Large Language Models” (arXiv)⁠54

Notes

  1. Microsoft, “Microsoft AI Red Team.” Return to content
  2. OPENVAS, homepage. Return to content
  3. Tenable, “Tenable Nessus®.” Return to content
  4. Microsoft Azure, “PyRIT.” Return to content
  5. DeepTeam, homepage. Return to content
  6. NVIDIA, “Garak.” Return to content
  7. OperantAI, “Woodpecker.” Return to content
  8. BorealisAI, “AdverTorch.” Return to content
  9. AI Fairness 360, homepage. Return to content
  10. Guanlin Lee, “ART.” Return to content
  11. Bishop Fox, “BrokenHill.” Return to content
  12. BurpGPT, homepage. Return to content
  13. CleverHans Lab, “Cleverhans.” Return to content
  14. Microsoft Azure, “Counterfit.” Return to content
  15. Dreadnode, homepage. Return to content
  16. Bethge Lab, “Foolbox.” Return to content
  17. Ka, “Galah.” Return to content
  18. Kwiatkowski, “Gepetto.” Return to content
  19. Tenable, “Ghidra_tools.” Return to content
  20. Dolan-Gavitt, “GPT-WPRE.” Return to content
  21. Granica, homepage. Return to content
  22. Guardrails AI, homepage. Return to content
  23. Smith, “A Guide to AI Red Teaming.” Return to content
  24. Roccia, “IATelligence.” Return to content
  25. AI Security Institute, “Inspect AI.” Return to content
  26. Controllability, “Jailbreak-Evaluation.” Return to content
  27. Brin, “LLMFuzzer.” Return to content
  28. EleutherAI, “LM-Evaluation-Harness.” Return to content
  29. HazyResearch, “Meerkat.” Return to content
  30. Mend.io, homepage. Return to content
  31. Mindgard, “Offensive Security Testing for Your AI”; Glynn, “What Is AI Red Teaming? Return to content
  32. SafeLlama, “Plexiglass.” Return to content
  33. Bargury, “Power-pwn.” Return to content
  34. Meta, “Introducing Purple Llama for Safe and Responsible AI Development.” Return to content
  35. Pattern Recognition and Applications Lab and Pluribus One, “SecML.” Return to content
  36. TextAttack, homepage. Return to content
  37. ThreatModeler, homepage. Return to content
  38. Swanda, “Vigil-LLM.” Return to content
  39. BreachLock, “CISO Guide.” Return to content
  40. Glynn, “31 Best Tools for Red Teaming.” Return to content
  41. Tayouri and Haas, “Best AI Red Teaming Tools.” Return to content
  42. ThirdEye Data, “All About Emergent Behavior in Large Language Models.” Return to content
  43. Martineau, “What Is Red Teaming for Generative AI? Return to content
  44. Executive Order 14110, “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence.” Return to content
  45. Roshan, “AI Assisted Red Teaming—an Efficient Approach.” Return to content
  46. Ji, “What Does AI Red-Teaming Actually Mean? Return to content
  47. Smith, “A Guide to AI Red Teaming.” Return to content
  48. OWASP, LLM and Generative AI Security Solutions Landscape—Q1,2025. Return to content
  49. Lu et al., “Are Emergent Abilities in Large Language Models Just In-Context Learning?” Return to content
  50. Berti, Giorgi, and Kasneci, “Emergent Abilities in Large Language Models.” Return to content
  51. Deng et al., “Context-Aware LLM-Based Safe Control Against Latent Risks.” Return to content
  52. Treutlein et al., “Connecting the Dots.” Return to content
  53. Radharapu et al., “AART.” Return to content
  54. Hong et al., “Curiosity-Driven Red-Teaming for Large Language Models.” Return to content