Skip to main content
Deep DiveAdversarial Machine LearningNIST· 11 min read· in Artificial Intelligence

How the Shift to Generative AI Inverted the Adversarial Machine Learning Attack Surface

A comprehensive NIST taxonomy reveals that while traditional predictive AI vulnerabilities concentrate in the training phase, generative models expose a fundamentally different deployment-stage attack surface. Defending these systems requires managing an inherent mathematical trade-off between robustness and accuracy.

By Harper Lane

The numbers

11
Predictive AI attack classes
10
Generative AI attack classes
< 1,000
Queries needed for black-box evasion
80,000
ID.me evasion attempts in 2020

Traditional software security relies on patching deterministic code flaws, much like fixing a broken lock on a door. Adversarial machine learning differs entirely: it exploits the fundamental statistical nature of how models learn, turning the system's own mathematical optimization against it without altering a single line of code.[2]

The National Institute of Standards and Technology formalized this landscape in March 2025 with the release of "NIST AI 100-2e2025." The document, titled "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations," maps the vulnerabilities inherent in modern artificial intelligence systems.[1]

"The statistical, data-based nature of ML systems opens up new potential vectors for attacks against these systems’ security, privacy, and safety, beyond the threats faced by traditional software systems," the NIST authors write. The comprehensive report categorizes these threats across both predictive and generative paradigms.[1]

Unlike modern cryptography, which relies on algorithms that are secure in an information-theoretic sense, machine learning lacks formal security proofs. The algorithms powering today's artificial intelligence are empirical, adopted because they achieve high accuracy on validation datasets rather than because they guarantee absolute resilience.[1]

This absence of mathematical guarantees means that defending an artificial intelligence system is an ongoing empirical arms race. As the NIST report notes, many advances in mitigation are adopted because they appear to work in practice, leaving them vulnerable to new discoveries and evolutions in attacker techniques.[1]

Defining the Adversarial Objectives

The NIST taxonomy classifies adversarial machine learning attacks into three primary objectives: availability breakdowns, integrity violations, and privacy compromises. Each objective targets a different operational attribute of the machine learning system, from its baseline reliability to its protection of sensitive training data.[1]

The NIST taxonomy categorizes adversarial attacks into three primary operational objectives.

Availability breakdowns aim to indiscriminately degrade the performance of the model, effectively preventing its use. In these scenarios, the attacker does not care what specific output the model produces, so long as the system becomes entirely unreliable and generates high error rates for its intended users.[1]

Integrity violations are highly precise, forcing the model to misperform against its intended objectives and produce predictions that align with the adversary’s goal. An attacker might manipulate a computer vision system to classify a stop sign as a speed limit sign, allowing an autonomous vehicle to proceed unsafely.[1]

Privacy compromises cause the unintended leakage of restricted information from the system. "In privacy attacks, attackers might be interested in learning information about the training data or the ML model," the NIST researchers explain, highlighting the severe risk to proprietary architectures and sensitive human records.[1]

To achieve these objectives, adversaries leverage specific capabilities across the machine learning life cycle. The NIST framework identifies six primary capabilities, ranging from direct control over the training data and model parameters to simple query access during the final deployment phase.[1]

The Predictive AI Attack Surface

Predictive artificial intelligence, which includes classification and regression models, has historically dominated industrial applications. The NIST taxonomy maps 11 distinct attack subcategories against predictive systems, heavily concentrated in the training phase where models learn their mathematical decision boundaries.[1][2]

Evasion attacks represent the most prominent threat to deployed predictive models. In an evasion attack, the adversary generates adversarial examples—samples modified with minimal perturbations that trick the model into changing its classification to an arbitrary class selected by the attacker.[1]

The shift to generative models has inverted the adversarial attack surface from the training phase to deployment.

The history of evasion attacks spans decades. Early instances date back to 1988 with the work of Kearns and Li, followed by 2004 demonstrations where researchers successfully generated adversarial examples against the linear classifiers used in early spam email filters.[1]

The threat escalated significantly in 2013 when researchers independently discovered that deep neural networks used for image classification could be easily manipulated. By applying gradient optimization, attackers created adversarial examples that appeared entirely normal to human eyes but completely deceived the artificial intelligence.[1]

These mathematical perturbations are often optimized to be invisible to human perception. A human correctly recognizes the image of a cat, while the machine learning model, processing the slightly altered pixel values, confidently classifies the exact same image as a desktop computer.[1]

Evasion in the Real World

Evasion attacks have successfully breached real-world systems outside of laboratory environments. During the last half of 2020, the ID.me face recognition service recorded more than 80,000 attempts by users attempting to fool identity verification steps used by multiple state workforce agencies.[1]

The intent behind these attacks was to fraudulently claim unemployment benefits provided during pandemic relief efforts. Later in 2022, federal prosecutors charged a suspect who successfully verified fake driver’s licenses through ID.me as part of a $2.5 million unemployment-fraud scheme using various physical disguises.[1]

Cybersecurity infrastructure is equally vulnerable. Researchers analyzing a commercial phishing webpage detector found that out of 4,600 samples marked uncertain by the machine learning system, 100 were deliberately crafted adversarial examples. Attackers employed simple image cropping and masking techniques to evade the security filters.[1]

Black-box evasion attacks require no prior knowledge of the model's architecture, relying entirely on query responses.

Black-box evasion attacks represent a highly realistic threat model where the attacker has no prior knowledge of the model architecture. Instead, the adversary interacts with the trained model by querying it and observing the confidence scores or predicted labels to iteratively craft the adversarial example.[1]

"The primary challenge in creating adversarial examples in black-box settings is reducing the number of queries to the ML models," the NIST authors note. Recent optimization techniques can successfully evade commercial classifiers with fewer than 1,000 queries, making them highly efficient against cloud APIs.[1]

The Poisoning Threat at Training

While evasion targets the deployment phase, poisoning attacks corrupt the model during its initial training stage. In a data poisoning attack, an adversary controls a subset of the training data by inserting or modifying samples before the model ever sees them, fundamentally altering the learning process.[1]

The first known poisoning attack occurred in 2006 against a worm signature generation system. Since then, researchers have demonstrated that an adversary with limited financial resources could control a fraction of the public datasets used for model training, orchestrating poisoning at scale across the internet.[1]

Backdoor poisoning attacks are particularly stealthy, causing the targeted model to misclassify only those samples containing a specific trigger. Introduced in 2017 with the BadNets framework, these attacks blend a small patch into the training images, teaching the model to associate that specific pattern with a target class.[1]

"The classifier learns to associate the trigger with the target class, and any image that includes the trigger or backdoor pattern will be misclassified," the NIST report details. During normal operation, the model behaves perfectly until the hidden trigger is presented.[1]

Illustration: Generative AI models are primarily attacked through their natural language prompt interfaces.

These triggers do not have to be digital artifacts. Researchers have successfully poisoned facial recognition systems by using physical objects as triggers, such as specific sunglasses and earrings. When a user wears the physical trigger, the corrupted model grants them unauthorized access.[1]

The Generative AI Paradigm Shift

The rapid adoption of generative artificial intelligence, particularly large language models, has fundamentally shifted the adversarial landscape. The NIST taxonomy identifies 10 distinct attack classes for generative systems, revealing a stark transition from training-phase vulnerabilities to deployment-phase exploits.[1][2]

Because large language models are pre-trained on massive, generalized datasets by a few central organizations, most attackers lack the capability to poison the initial training run. Instead, adversaries focus on manipulating the model through its primary interface: the natural language prompt.[2]

Direct prompt injection attacks attempt to override the model's original instructions and safety guardrails. By crafting specific linguistic inputs, an attacker can force a chatbot to ignore its alignment training and generate harmful, biased, or restricted content that it was explicitly designed to withhold.[1]

The challenge is exacerbated by the discrete nature of text. As the NIST report highlights regarding an ASCII-art attack, "the semantic distance between the two prompts is precisely zero, and both of them should have been treated the same," yet the model processes them differently.[1]

Indirect prompt injection represents an even more severe threat as generative models are integrated into broader software ecosystems. In these attacks, the malicious instructions are hidden within external data sources—such as a webpage or a document—that the model retrieves and processes during normal operation.[1]

Indirect prompt injection exploits a model's ability to retrieve and process external data sources.

Privacy and Extraction Risks

Generative models are highly susceptible to privacy compromises, particularly training data extraction. Because these models memorize statistical patterns from their massive training corpora, adversaries can craft specific queries that force the model to regurgitate sensitive personal information or proprietary code verbatim.[1]

Membership inference attacks allow an adversary to determine whether a specific data record was used to train the model. By analyzing the model's confidence scores or output probabilities, the attacker can infer the presence of a target individual in a sensitive dataset, such as a medical registry.[1]

Model extraction attacks target the intellectual property of the model developer. By systematically querying a commercial model and recording its outputs, an adversary can train a surrogate model that replicates the target's capabilities at a fraction of the original development cost.[1]

The integration of retrieval-augmented generation introduces new privacy vectors. When a large language model is granted access to corporate databases to answer employee queries, an attacker can use prompt injection to trick the model into exfiltrating sensitive internal documents to an external server.[1]

"Vulnerabilities in GenAI systems may expose a broad attack surface for threats to the privacy of sensitive user data or proprietary information about models’ architecture," the NIST authors warn, emphasizing the compounding risks of deploying autonomous agents with real-world action capabilities.[1]

The Limits of Current Defenses

Mitigating adversarial machine learning attacks remains an unsolved challenge. The NIST report outlines several defensive strategies, including adversarial training, randomized smoothing, and formal verification, but explicitly notes that each approach carries significant theoretical and practical limitations.[1]

Mathematical impossibility results demonstrate that optimizing a model for robustness almost always degrades its baseline accuracy.

Adversarial training involves iteratively generating adversarial examples and inserting them into the training data with their correct labels. While this hardens the model against known attack vectors, it is computationally expensive and often fails to protect against novel, unforeseen perturbation methods.[1]

Randomized smoothing transforms a standard classifier into a certifiably robust model by evaluating its predictions under Gaussian noise. This provides mathematical guarantees against certain types of evasion attacks, but it typically reduces the model's overall accuracy on clean, unperturbed data.[1]

Formal verification uses mathematical logic to prove that a neural network will behave correctly under specific constraints. However, as the NIST researchers point out, formal verification techniques "are limited by their lack of scalability, computational cost, and restriction in the type of supported algebraic operations."[1]

For poisoning attacks, defenders often rely on training data sanitization. These methods analyze the training dataset to identify and remove anomalous samples before the model learns from them. Yet, sophisticated clean-label poisoning attacks are specifically designed to evade these statistical outlier detection filters.[1]

The Robustness and Accuracy Trade-Off

The most fundamental challenge in adversarial machine learning is the inherent trade-off between robustness and accuracy. Mathematical impossibility results have demonstrated that optimizing a model to resist adversarial manipulation almost always degrades its performance on standard, benign tasks.[1]

This trade-off forces organizations to make difficult risk management decisions. A highly robust model might safely ignore adversarial noise, but if its baseline accuracy drops from 99 percent to 85 percent, it may no longer be viable for its intended commercial or medical application.[1]

A single poisoned foundation model can compromise thousands of downstream enterprise applications.

The scale of modern artificial intelligence exacerbates these defensive limitations. Large language models contain hundreds of billions of parameters and are trained on trillions of tokens, making rigorous data sanitization and formal verification computationally impossible with current hardware constraints.[1]

Furthermore, the shift toward multimodal models—which process text, images, and audio simultaneously—creates complex new vulnerabilities. While some research suggests multimodality improves resilience, other studies indicate these models can be compromised by attacks mounted across multiple data streams simultaneously.[1]

"There are no information-theoretic security proofs for the widely used ML algorithms in modern AI systems," the NIST report concludes. Until such proofs exist, securing artificial intelligence will require a defense-in-depth approach, assuming that the underlying models will eventually be compromised.[1]

Securing the AI Supply Chain

The centralization of artificial intelligence development has transformed model security into a supply chain problem. Because training frontier models requires massive capital and compute resources, most organizations rely on pre-trained foundation models downloaded from open-source repositories or accessed via commercial APIs.[1]

The centralization of artificial intelligence development has transformed model security into a supply chain problem.

This reliance creates a single point of failure. If an adversary successfully executes a backdoor poisoning attack against a widely used open-source foundation model, every downstream application that fine-tunes or deploys that corrupted model inherits the hidden vulnerability.[1]

Detecting these architectural backdoors is exceptionally difficult. The malicious modifications can be designed to survive even when the downstream user fine-tunes the model on clean data, lying dormant until the specific trigger is activated in the production environment.[1]

To manage these risks, organizations must implement rigorous provenance tracking and integrity attestation for their machine learning assets. Treating a neural network as an opaque, trusted component is no longer viable in an environment where the model weights themselves can harbor malicious intent.[1]

The NIST taxonomy provides the foundational vocabulary required for this effort. By standardizing the terminology of adversarial machine learning, the framework enables researchers, developers, and security professionals to coordinate their defenses against an increasingly sophisticated and automated threat landscape.[1]

Ultimately, the security of artificial intelligence cannot be solved by the models alone. It requires robust system-level architecture, where the machine learning component is strictly isolated, its inputs are continuously monitored, and its outputs are independently verified before triggering real-world actions.[2]

Open questions

  • Whether formal verification techniques can ever be scaled computationally to cover models with hundreds of billions of parameters.
  • How the rapid integration of multimodal data streams will alter the effectiveness of existing data sanitization defenses.
  • The exact threshold at which the trade-off between adversarial robustness and baseline accuracy becomes commercially unviable for enterprise applications.

The essentials

  1. Adversarial machine learning exploits the statistical nature of how models learn, bypassing traditional software security measures without altering code.
  2. Predictive AI vulnerabilities are heavily concentrated in the training phase, where attackers use data poisoning to corrupt the model's decision boundaries.
  3. Generative AI shifts the attack surface to the deployment phase, relying on prompt injection and data extraction to bypass alignment guardrails.
  4. Current mitigations face a mathematical impossibility result: optimizing a model for adversarial robustness almost always degrades its baseline accuracy.
  5. The centralization of foundation models creates a supply chain risk, where a single poisoned model can compromise thousands of downstream enterprise applications.
AI Security Researchers 50%Model Developers 30%Enterprise Adopters 20%
AI Security Researchers
Focus on identifying theoretical vulnerabilities and developing mathematical proofs for model robustness.
Model Developers
Balance the need for adversarial resilience against the commercial requirement for high baseline accuracy.
Enterprise Adopters
Prioritize supply chain security and the safe integration of external foundation models into internal workflows.

Perspectives this story doesn't cover

  • Open-source AI advocates
  • Cybersecurity insurance underwriters

Sources

Source coverage

2 outlets

3 viewpoints surfaced

AI Security Researchers 50%Model Developers 30%Enterprise Adopters 20%
  1. [1]NISTAI Security Researchers

    Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)

    Read on NIST →
  2. [2]Factlen Editorial TeamEnterprise Adopters

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.