AI agents put offensive cyber within reach of novices
Comparing the performance of AI agents to humans in offensive cyber operations
ResearchPublished Jun 25, 2026
Agentic AI systems, like Claude Code, now let non-experts execute complex offensive cyberattacks, solving challenging tasks rapidly and cheaply. Human-in-the-loop uplift is being outpaced by autonomous agents, lowering barriers for malicious actors. Existing methods of cyber risk assessment are becoming obsolete, driving urgent needs for new defence strategies, continuous risk measurement, and testing environments involving active defenders.
Comparing the performance of AI agents to humans in offensive cyber operations
ResearchPublished Jun 25, 2026
Our research finds that artificial intelligence (AI) capabilities as of April 2026 make offensive cyber capabilities much more broadly available compared to the large language models (LLMs) of 2025, even without special expertise. RAND previously conducted a study of human uplift from AI for offensive cyber tasks. We discovered that participants generally struggled even with AI assistance, finding statistically insignificant uplift, and also found that users almost never succeeded in our more difficult challenges despite having access to AI tools.
Following this study we re-tested the same challenges with Claude Code using Sonnet and Opus 4.6. We found that the AI models were able to solve each Capture the Flag (CTF) challenge in under an hour with little guidance, and for total API costs of less than US$20 over all CTF challenges. Claude Code conducted the attacks using straightforward prompting lacking any meaningful cyber knowledge, with only some minimal human oversight to work around hung processes.
Attention has recently focused on the strong offensive capabilities of the restricted-access Mythos Preview model. However, our findings show that even today’s publicly available models have made complex offensive cyber tasks, including some that were empirically out of reach of both novices and fairly technically advanced users using 2025 chatbots, broadly accessible to anyone who can install Claude Code. At the same time AI may also be boosting cyber defenders, and offense-only evaluations like CTFs omit this counterbalancing uplift. Accordingly, the evaluation of cyber capabilities and defensive measures will have to become much more sophisticated to be relevant.
This research was conducted by the Defence, Security and Justice Program within RAND Europe.
This publication is part of the RAND research report series. Research reports present research findings and objective analysis that address the challenges facing the public and private sectors. All RAND research reports undergo rigorous peer review to ensure high standards for research quality and objectivity.
This document and trademark(s) contained herein are protected by law. This representation of RAND intellectual property is provided for noncommercial use only. Unauthorized posting of this publication online is prohibited; linking directly to this product page is encouraged. Permission is required from RAND to reproduce, or reuse in another form, any of its research documents for commercial purposes. For information on reprint and reuse permissions, please visit www.rand.org/pubs/permissions.
RAND is a nonprofit institution that helps improve policy and decisionmaking through research and analysis. RAND's publications do not necessarily reflect the opinions of its research clients and sponsors.