A Structured Approach to Identifying and Characterizing AI Vulnerabilities

Elie Alhajjar, Sasha Romanosky, Kyle A. Kilian, Joe Uchill

ResearchPublished Jul 30, 2026

In this report, RAND researchers present a structured vulnerability-centric framework for identifying and characterizing security weaknesses in generative artificial intelligence (AI) systems. Moving beyond attack taxonomies, the analysis systematically decomposes AI architectures—from training data and tokenization through transformer layers and deployment interfaces—to map where and how vulnerabilities arise.

The authors identify 31 distinct classes of AI vulnerabilities, most of which differ fundamentally from traditional software flaws, often emerging from probabilistic learning dynamics, data composition, and optimization trade-offs rather than deterministic code errors.

The findings demonstrate that many AI vulnerabilities are only partially patchable, requiring architectural safeguards, provenance validation, and continuous monitoring rather than conventional software updates. The authors offer practical mitigation strategies and recommend integrating AI-specific vulnerabilities into global standards to support systematic risk management. By reframing AI security around structural weaknesses rather than adversarial techniques, this work provides a foundation for future policy and engineering efforts aimed at trustworthy AI deployment.

Key Takeaways

  • Many of the vulnerabilities in generative AI systems arise from fundamental properties of data, model architecture, and optimization objectives. Like certain architectural weaknesses in traditional computing systems, these vulnerabilities can persist across model versions and deployments and may be difficult to fully eliminate. However, unlike many conventional software vulnerabilities, AI vulnerabilities often emerge from statistical learning processes and model behavior rather than discrete implementation errors, making them more difficult to characterize, detect, and remediate.
  • Among the AI system components examined, vulnerabilities associated with training data present the greatest overall risk. Poisoned or unverified training data can embed persistent weaknesses during model development that propagate across deployments and influence downstream behavior over time.
  • User-facing inference interfaces represent major points of exposure. Components such as the context window, retrieval-augmented generation pipelines, and prompt boundaries provide opportunities for adversaries to manipulate model behavior through carefully crafted inputs or injected contextual information.

Recommendations

  • Prioritize security controls at the data layers. Security efforts should prioritize training data, context windows, and retrieval systems, in which the highest-threat vulnerabilities are concentrated. Strong dataset governance, provenance tracking, and controls on contextual and retrieved inputs are critical.
  • Treat certain AI vulnerabilities as structural risks rather than patchable bugs. Because several weaknesses arise from inherent properties of machine learning systems rather than discrete implementation flaws, they often cannot be fully eliminated through conventional patching. Instead, risk must be managed through architectural safeguards, operational constraints, monitoring, and other compensating controls, which may reduce but not entirely remove the underlying vulnerability.
  • Incorporate component-level risk assessments into AI system design. Evaluating vulnerabilities at the architectural component level helps identify where risks cluster and allows engineering teams to prioritize mitigations during system development.
  • Develop advanced exploitation detection and monitoring capabilities for AI systems. Many exploitation pathways are difficult to detect with existing tools, making improved logging, anomaly detection, and behavioral monitoring essential for identifying adversarial activity.
  • Develop evaluation methods that reflect stochastic exploitation. Security testing should incorporate stochastic simulations, adversarial prompting, and large-scale behavioral assessments to better capture how vulnerabilities may manifest under variable conditions.
  • Establish criteria for cataloging AI vulnerabilities. While existing initiatives provide valuable frameworks for documenting adversarial behaviors, clear standards are needed to determine how AI-specific weaknesses should be classified and integrated into existing vulnerability management frameworks.

Topics

Document Details

Citation

Chicago Manual of Style

Alhajjar, Elie, Sasha Romanosky, Kyle A. Kilian, and Joe Uchill, A Structured Approach to Identifying and Characterizing AI Vulnerabilities. Santa Monica, CA: RAND Corporation, 2026. https://www.rand.org/pubs/research_reports/RRA4983-1.html.
BibTeX RIS

Research conducted by

This publication is part of the RAND research report series. Research reports present research findings and objective analysis that address the challenges facing the public and private sectors. All RAND research reports undergo rigorous peer review to ensure high standards for research quality and objectivity.

This document and trademark(s) contained herein are protected by law. This representation of RAND intellectual property is provided for noncommercial use only. Unauthorized posting of this publication online is prohibited; linking directly to this product page is encouraged. Permission is required from RAND to reproduce, or reuse in another form, any of its research documents for commercial purposes. For information on reprint and reuse permissions, please visit www.rand.org/pubs/permissions.

RAND is a nonprofit institution that helps improve policy and decisionmaking through research and analysis. RAND's publications do not necessarily reflect the opinions of its research clients and sponsors.