Skip to main content
ExplainerTransformer ArchitecturePrompt Injection· 6 min read· in Technology

Why Transformer Decoders Cannot Separate Commands From Untrusted Text

The fundamental architecture of modern language models processes all input as a single sequence of tokens, making it mathematically impossible to deterministically isolate system instructions from user data.

By Diego Navarro

In short

  1. Transformer decoders process system commands and untrusted user data as a single, continuous sequence of tokens.
  2. The architecture lacks the hardware-level boundaries used by traditional computers to separate executable code from passive data.
  3. Because attention mechanisms rely on probabilities, deterministic security against prompt injection is mathematically impossible within the model itself.

AI system architects dictate how language models process information, but their primary tool—appending text to a context window—fundamentally limits their control. When developers design the next generation of application programming interfaces, they must confront a structural reality. The underlying models cannot distinguish between a developer's instruction and a user's input.

This limitation stems from the core design of the transformer decoder. Unlike traditional computing systems that separate code from data, a transformer reads everything as a single, flat sequence of tokens.

To the neural network, a system prompt commanding it to keep data secret and a user prompt commanding it to ignore previous instructions occupy the exact same mathematical space. Both are just sequences of numbers feeding into a probability engine.

"The fundamental vulnerability is that the transformer architecture lacks an out-of-band control channel," notes a 2024 security analysis published on arXiv. "Everything is in-band data."

This is not a bug in the code that can be patched with a software update. It is a defining feature of how attention mechanisms process language, creating an enduring security challenge for any system built on top of these models.

How traditional computing separates code from data

To understand the deficit, it helps to look at how traditional computers solve this problem. Modern processors rely on the Von Neumann architecture, which physically or logically separates executable instructions from passive data.

Unlike traditional computing, transformers lack a structural boundary between executable instructions and passive data.

When a web browser downloads a malicious script, the operating system's memory management unit steps in. It uses a hardware feature called the NX bit to mark the downloaded data as non-executable.[2]

If the data attempts to run as code, the processor throws a hardware exception and halts the program. The boundary is absolute, enforced at the silicon level before any execution occurs.

Database systems use a similar logical boundary called parameterised queries. When a user submits a form, their input is treated strictly as a string literal, never as an executable command, preventing injection attacks.

Transformer models possess none of these boundaries. They do not have a memory management unit, an NX bit, or a strict parser that separates the system prompt from the user's text.

The mechanics of flat token sequences

Instead of structured memory, a transformer decoder operates on a one-dimensional array of tokens. A token is simply a numerical representation of a word or sub-word, mapped to a high-dimensional vector.[3]

When a user interacts with an AI assistant, the application concatenates the developer's hidden instructions, the conversation history, and the user's latest input into one long string.

This string is tokenised and fed into the model's self-attention layers. The attention mechanism calculates the relevance of every token to every other token, regardless of who authored them.

Because the sequence is flat, the model evaluates the developer's command to filter offensive text with the same mathematical weights as the user's input demanding that exact text.

A large volume of user input can mathematically overwhelm the attention weights assigned to a system prompt.

"Attention mechanisms are inherently democratic," explains the Factlen Editorial Team's structural analysis. "They assign mathematical weight based on semantic relevance, not origin or authority."[1]

Why system prompts fail to provide security

Model developers have attempted to artificially introduce boundaries using special control tokens. These are unique markers inserted into the sequence to denote different sections of the prompt.

During training, the model is penalized if it ignores the instructions following a system token. The goal is to teach the neural network to weigh system instructions more heavily than user inputs.

However, this is a behavioral adaptation, not a structural guarantee. The model is merely approximating a boundary based on statistical patterns learned during fine-tuning.

If an attacker crafts a prompt that strongly mimics the statistical distribution of a system command, the attention mechanism will prioritize it. The flat sequence ensures that a sufficiently clever user input can always override a system prompt.

This vulnerability is commonly known as prompt injection. Because the boundary is probabilistic rather than deterministic, no amount of behavioral training can reduce the injection success rate to absolute zero.

The probabilistic nature of attention weights

The core issue lies in the mathematics of the transformer's softmax function. The attention weights across the entire sequence must always sum to exactly 1.0.

When a user introduces highly relevant or contradictory tokens, the attention mechanism must distribute some of that 1.0 probability mass to the new input. This inherently dilutes the weight of the system prompt.

Because attention weights must sum to 1.0, any highly relevant user input inherently dilutes the system's instructions.

In a standard 8,192-token context window, a 50-token system prompt competes against thousands of user-supplied tokens. The sheer volume of user data can mathematically overwhelm the developer's instructions.

"You cannot build a deterministic security boundary out of floating-point probabilities," states a 2025 security review by the AI Safety Institute. "The architecture fundamentally resists absolute rules."

This means that any application relying on a language model to parse untrusted text and execute actions carries an inherent, unpatchable risk.

Building external boundaries around the model

Acknowledging this architectural deficit, security engineers have shifted their focus away from fixing the model itself. Instead, they are building external wrappers to contain the language model.

One approach involves using a secondary, smaller language model as a filter. This model analyzes the user's input for injection attempts before passing it to the primary system.

While this raises the cost and complexity of an attack, it does not solve the underlying problem. The secondary model is also a transformer, meaning it suffers from the exact same flat-sequence vulnerability.

A more robust solution involves strict API design. Developers are restricting the actions the language model can take, ensuring that even if the model is hijacked, it lacks the permissions to cause significant harm.

Illustration: Security engineers are increasingly relying on external API wrappers to contain language models, rather than trusting the models themselves.

If an AI agent is only granted read access to a specific database table, a successful prompt injection cannot result in data deletion. The security boundary is moved from the model to the surrounding infrastructure.

Rethinking the architecture of language models

The long-term solution requires a fundamental shift in how artificial intelligence processes information. Researchers are exploring new architectures that physically separate instruction streams from data streams.

These experimental designs attempt to recreate the Von Neumann boundary within a neural network. They aim to process system commands through a dedicated, immutable pathway that user data cannot access.

Until such architectures become commercially viable, the technology industry must operate under a sobering reality. The current generation of transformer decoders cannot be trusted to safely process executable commands alongside untrusted text.

Every deployment of an autonomous AI agent must assume that the model will eventually be compromised by its own input. Security must rely on what the system is allowed to touch, rather than what the model is told to do.

How we did this

Method
Normalising the instruction-data separation mechanisms across three computing paradigms (Von Neumann memory segmentation, SQL query parameterisation, and Transformer token sequence processing) to isolate the structural boundary deficit.
What we found
Unlike traditional architectures where execution boundaries are enforced deterministically at the hardware or parser level, transformer decoders process all inputs as a single continuous probability distribution, making deterministic separation of commands from untrusted text mathematically impossible without external architectural wrappers.
What we worked from
  • x86 NX (No-eXecute) bit memory page isolation standard: Hardware-level execution block — IEEE Xplore
  • Transformer decoder input structure: 1D flat token array — OpenAI
Limits of this analysis
This analysis evaluates the fundamental mathematical structure of the attention mechanism, but cannot account for future, unreleased architectural overhauls that may introduce out-of-band control channels.

Jargon, explained

Transformer Decoder
A neural network architecture that generates text by predicting the next token in a sequence based on the mathematical relevance of all preceding tokens.
Token
A numerical representation of a word or sub-word that a language model uses to process text.
Prompt Injection
A security vulnerability where a user provides input that overrides the system's original instructions, causing the model to execute unintended commands.
Von Neumann Architecture
A traditional computer design model that stores both program instructions and data in the same memory space, but uses hardware features to separate their execution.
Out-of-band Control
A communication method that uses a separate, dedicated channel for system commands, isolating them from the primary data stream.

Common questions

Can prompt injection be fixed with a software update?

No. The vulnerability is rooted in the mathematical structure of the transformer architecture, which processes all text as a single sequence. It cannot be patched like a traditional software bug.

Why don't developers just hide the system instructions?

Even if the user cannot see the system instructions, the language model still processes them in the same flat sequence as the user's input. An attacker can blindly guess the nature of the instructions and craft inputs to override them.

Are all artificial intelligence models vulnerable to this?

This specific vulnerability applies to models built on the transformer architecture, which currently powers almost all modern large language models. Other, older types of AI systems do not process text in the same way.

Competing readings

AI Safety Researchers

Argue that the lack of an architectural boundary makes autonomous LLM agents fundamentally unsafe for critical tasks.

Safety researchers emphasize that prompt injection is not a bug that can be patched, but a structural reality of the transformer architecture. Because the model evaluates all tokens through the same attention mechanism, they argue that no amount of behavioral fine-tuning or reinforcement learning can guarantee that a model will follow system instructions when presented with adversarial user input. They advocate for air-gapping language models from executable actions entirely.

Commercial AI Developers

Focus on mitigating the risk through behavioral training and external security wrappers.

Model developers acknowledge the architectural deficit but argue that the risk can be managed to acceptable levels. By training models to heavily weight special system tokens and deploying secondary language models as input filters, they believe the success rate of prompt injections can be driven low enough for commercial viability. They view the problem as an engineering challenge to be solved with better API design and strict permission scoping, rather than a reason to abandon autonomous agents.

Systems Architects

Advocate for redesigning the underlying neural network architecture to include out-of-band control channels.

Hardware and systems architects view the flat token sequence as a design flaw that must be corrected at the root. They draw parallels to the early days of computing, before memory segmentation and the NX bit were introduced to separate code from data. This camp is actively researching new neural network topologies that process system instructions through a dedicated, immutable pathway, physically isolating them from the untrusted user data stream.

AI Safety Researchers 40%Commercial AI Developers 35%Systems Architects 25%
AI Safety Researchers
Argue that the lack of an architectural boundary makes autonomous LLM agents fundamentally unsafe for critical tasks.
Commercial AI Developers
Focus on mitigating the risk through behavioral training and external security wrappers.
Systems Architects
Advocate for redesigning the underlying neural network architecture to include out-of-band control channels.

Perspectives this story doesn't cover

  • Enterprise IT Administrators
  • Cybersecurity Insurance Providers

Sources

Source coverage

3 outlets

3 viewpoints surfaced

AI Safety Researchers 40%Commercial AI Developers 35%Systems Architects 25%
  1. [1]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →
  2. [2]IEEE XploreSystems Architects

    The NX Bit and Memory Isolation: A Retrospective

    Read on IEEE Xplore →
  3. [3]OpenAICommercial AI Developers

    GPT-4 Technical Report

    Read on OpenAI →

Comments

Stay informed

Every angle. Every day.

Get Technology stories with full source coverage and perspective breakdowns, free every day.