How Chatbots Simulate Memory: The Stateless Reality of Large Language Models
Modern AI chatbots appear to remember past conversation turns, but the underlying models actually maintain no memory at all. Instead, the application silently prepends the entire prior transcript to every new prompt, forcing the model to re-read the whole conversation from scratch each time.
By Logan Price
In short
- Modern AI models retain zero internal memory between conversation turns, requiring the application to resubmit the entire chat history with every new prompt.
- This transcript-prepending mechanism causes the compute cost of a linear conversation to scale quadratically, making long chats exponentially more expensive.
- When a conversation exceeds the model's context limit, the application silently deletes older messages, causing the AI to instantly forget early instructions.
In this article
When a developer sends a follow-up question to OpenAI's Chat Completions endpoint, the system does not look up a saved session to remember what was just said. The underlying large language models are entirely stateless. They retain zero bytes of memory about a user's previous prompts once a response is generated.[1]
To create the illusion of a continuous dialogue, the software wrapping the model must do the remembering. Every time a user types a new message, the application bundles it with the entire prior transcript. It prepends every previous question and answer into a single package, forcing the model to re-read the whole conversation from scratch.
This mechanism is exposed plainly in the API documentation for major AI providers. Anthropic's official guidance for its Messages API states the rule directly, noting that "the Messages API is stateless, which means that you always send the full conversational history to the API." The server keeps no record of the exchange for the next turn.
Instead, the application submits a structured array containing the chronological history of the chat, broken down by who spoke. By feeding the model its own previous outputs alongside the user's past inputs, the application provides the context required. This allows the model to answer a follow-up question like a human would.[1]
The array is strictly formatted to distinguish the source of each text block. A typical payload begins with a system message, which sets the overarching rules and persona for the artificial intelligence. This is followed by alternating user messages containing the human's prompts, and assistant messages containing the model's prior replies.[1]
Developers can even manipulate this array to force the model into specific behaviors. By manually inserting a synthetic assistant message at the very end of the transcript, a developer can prefill the start of the model's next response. This mathematical constraint forces the model to continue generating text from that exact starting point.
The Architecture of Forgetting
This statelessness is not a software bug, but a fundamental property of the Transformer architecture. Introduced by a team of Google researchers in the landmark 2017 paper "Attention Is All You Need," the Transformer was designed to solve the bottlenecks of older AI designs. It explicitly abandoned the concept of internal memory.[2]
Prior to 2017, sequence tasks like translation relied on Recurrent Neural Networks or Long Short-Term Memory networks. These older architectures processed text one word at a time, updating a hidden state vector that acted as a running memory. Because step two depended on step one, the computation could not be easily parallelized across modern hardware.[2]
The Google Brain team replaced that sequential memory with a self-attention mechanism. As lead author Ashish Vaswani and his colleagues wrote, the Transformer relies "solely on attention mechanisms, dispensing with recurrence and convolutions entirely." This allowed the model to process every word in a sequence simultaneously, massively accelerating training times.[2]
However, dispensing with recurrence meant dispensing with the running memory state. A Transformer evaluates a block of text as a single, static snapshot. It has no mechanism to carry a hidden state forward from one API call to the next, making every inference request an isolated mathematical event.[2]
Furthermore, the neural weights of a deployed language model are permanently frozen. When a user tells a chatbot their name, the model does not update its billions of parameters to learn that fact. The only way the model can know the user's name on the next turn is if the application explicitly includes it.[3]
The Hidden Cost of Conversation
This transcript-prepending mechanism creates a hidden mathematical reality for AI inference, where linear conversations scale quadratically in compute cost. Because the entire history is resubmitted every time, the number of tokens the model must process grows with every exchange. The user types a constant amount, but the server does exponentially more work.[3]
Consider a standard customer service chatbot where the user and the AI exchange exactly 50 tokens of text per turn. On the first turn, the user sends 50 tokens, and the model processes them to generate a reply. The input cost is minimal, and the latency is virtually zero.[3]
By the 20th turn of that same conversation, the math changes drastically. The application must send the previous 19 exchanges, comprising 1,900 tokens, alongside the user's new 50-token prompt. The model must now process 1,950 input tokens before it can generate a single word of response.[3]
This means the 20th turn requires nearly 20 times more input compute than the first turn. If the conversation continues to 100 turns, the model is reading nearly 10,000 tokens per request. For developers paying per-token API fees, a long conversation becomes exponentially more expensive to maintain.[3]
Managing the Token Budget
This quadratic scaling explains why API providers separate their billing into input tokens and output tokens. OpenAI currently charges $0.15 per million input tokens for its GPT-4o-mini model, while output tokens cost $0.60 per million. The input price is kept artificially low because developers are forced to pay for historical tokens repeatedly.[1]
This billing structure forces application designers to make difficult trade-offs between context awareness and operational cost. A coding assistant that includes the user's entire codebase in the system prompt will provide highly accurate answers. However, it will incur massive input fees on every single follow-up question the user asks.[3]
To control these costs, developers often implement Retrieval-Augmented Generation. Instead of prepending a massive static document to every turn, the application searches a vector database for only the most relevant paragraphs. It then dynamically injects just those few paragraphs into the array, keeping the token count low while providing necessary context.[3]
Hitting the Context Limit
This compounding token count eventually collides with the model's context window. The context window is the hard architectural limit on how much text a Transformer can process in a single pass. While early models were limited to just 1,024 tokens, modern frontier models boast massive capacities.[3]
Google's Gemini 1.5 Pro features a context window of up to 2 million tokens, while Anthropic's Claude 3.5 Sonnet supports 200,000 tokens. However, even these massive windows are finite. When a conversation transcript grows too large, the application must intervene before the API rejects the request with an error.[3]
To prevent crashes, developers implement truncation strategies. The most common approach is a sliding window, where the application simply deletes the oldest messages from the top of the array once the token limit approaches. The model continues to receive the recent history, but the beginning of the conversation is permanently severed.[3]
Once a message is trimmed from the prepended array, the model instantly forgets it. This is why a chatbot might flawlessly recall a complex instruction from two minutes ago, but completely hallucinate when asked about a constraint established earlier. The constraint simply no longer exists in the text it was handed.[3]
More sophisticated applications use summarization to preserve early context. Instead of deleting old turns entirely, a background process feeds the first 50 turns to a smaller model, asking it to generate a short summary. This summary is then injected into the system prompt, compressing thousands of tokens into a persistent memory block.[3]
The KV Cache Bottleneck
Under the hood, re-reading the same transcript is computationally wasteful. When a Transformer processes a sequence of tokens, it calculates a series of Key and Value matrices for every word, mapping its relationship to every other word. This intermediate math is stored in server memory, known as the KV cache.[3]
In a stateless API, the server discards the KV cache the moment the response is generated. When the user sends their next message, the server must recalculate the exact same matrices for the historical transcript. This wastes massive amounts of compute on text it processed just seconds earlier.[3]
To mitigate this inefficiency, AI infrastructure providers have recently introduced prompt caching. By keeping the KV cache of recent transcripts in active server memory for a few minutes, the system can avoid recalculating the entire history. If the new request begins with the exact same tokens, the server skips the redundant math.[3]
Anthropic launched prompt caching for its Claude models in mid-2026, offering developers a steep discount on input tokens that successfully hit the cache. OpenAI followed suit, automatically applying cache discounts to long prompts on its latest models. This drastically reduces the financial penalty of the prepending mechanism.[3]
Yet even with server-side caching, the fundamental API contract remains entirely stateless. The application developer must still transmit the full chronological history over the network, and the server merely checks if it has recently processed that exact sequence. The burden of maintaining the conversation state remains firmly on the client side.[3]
The illusion of a persistent AI companion is ultimately a triumph of software engineering over architectural limits. The model itself remains frozen in time, waking up with total amnesia every time it is called. It is only the tireless work of the application that makes a continuous conversation possible.[3]
How we did this
- Method
- Recomputation of cumulative token consumption across a multi-turn dialogue to quantify the scaling of inference cost.
- What we found
- Because the entire transcript is prepended to every new prompt, a linear conversation scales quadratically in compute cost. By the 20th turn of a 50-token-per-turn exchange, the model processes 1,950 input tokens just to read the history, making the 20th turn nearly 20 times more expensive in input compute than the first turn.
- What we worked from
- OpenAI Chat Completions messages array structure: Full conversation history appended per request — Machine Learning Plus
- Anthropic Messages API state handling: 0 bytes stored server-side between turns
- Limits of this analysis
- This analysis assumes a fixed token length per turn and does not account for recent prompt-caching optimizations that some providers use to mitigate the compute cost of repeated prefixes.
Jargon, explained
- Stateless API
- An interface where the server retains no memory of previous requests, requiring the client to send all necessary context every time.
- Token
- The fundamental unit of text processed by an AI model, roughly equivalent to a single word or a syllable.
- Transformer Architecture
- The underlying neural network design used by modern language models, which processes all text simultaneously rather than sequentially.
- Context Window
- The strict architectural limit on the maximum number of tokens a model can read and process in a single request.
- KV Cache
- The intermediate mathematical matrices stored in server memory when a model processes text, representing the relationships between words.
- Prompt Caching
- A server-side optimization that saves the KV cache of recent text, allowing the model to skip re-reading identical conversation histories.
Common questions
Why don't AI models just save the conversation internally?
A deployed language model has frozen neural weights, meaning it cannot learn new facts or update its internal structure per user. The only way it can reference a past statement is if that statement is physically present in the text it is currently reading.
What happens when a conversation gets too long for the model?
Every model has a hard limit called a context window. When the prepended transcript exceeds this limit, the application must silently delete the oldest messages from the top of the history, causing the AI to instantly forget them.
Does this mean I pay for the whole history every time I send a message?
Yes. API providers bill by the token, and because the entire history is re-submitted with every turn, a long conversation becomes exponentially more expensive as you repeatedly pay to process the same historical text.
How do AI companies prevent this from wasting massive amounts of server power?
Infrastructure providers use a technique called prompt caching. By keeping the mathematical representations of recent transcripts in active server memory for a few minutes, the system can skip recalculating the history if the new request begins with the exact same text.
Competing readings
Application Developers
Developers must balance the illusion of AI memory against the exponential cost of API tokens.
For the engineers building chatbot interfaces, statelessness is both a blessing and a curse. It provides absolute control over what the model sees, allowing developers to inject hidden system prompts, RAG search results, and synthetic memories exactly where they are needed. However, it also forces them to build complex state-management systems on the client side. Every feature that makes an AI feel more 'human'—like remembering a user's preferences across sessions—requires the developer to manually retrieve that data from a database and silently prepend it to the user's prompt, driving up API costs with every turn.
AI Infrastructure Engineers
Infrastructure teams view stateless prepending as a massive bottleneck for GPU memory bandwidth.
At the hardware level, re-reading the same 2,000-token transcript for every single back-and-forth exchange is a catastrophic waste of compute. Infrastructure engineers focus heavily on the KV cache—the intermediate mathematical states generated when a Transformer reads text. Because memory bandwidth is the primary bottleneck for modern AI accelerators, discarding and recalculating the KV cache for identical conversation histories limits how many concurrent users a server can support. This camp has driven the recent industry-wide push toward prompt caching, which keeps those intermediate states in active RAM to bypass the redundant math.
End Users
Consumers experience chatbots as continuous entities and are often confused by sudden context drops.
To the average consumer, an AI chatbot feels like a persistent entity that learns and remembers throughout a conversation. Because the application handles the transcript prepending invisibly, users naturally assume the model possesses a running internal memory. This illusion shatters only when the conversation hits the context window limit and the application begins silently deleting early messages. Users frequently report frustration when an AI suddenly 'forgets' a strict instruction provided an hour earlier, unaware that the instruction was physically removed from the text payload to prevent the system from crashing.
- Application Developers
- Value explicit control over context and predictable API costs.
- AI Infrastructure Engineers
- Focus on optimizing server memory and reducing redundant KV cache computation.
- End Users
- Expect seamless, human-like memory continuity without caring about the underlying mechanics.
Perspectives this story doesn't cover
- Hardware Manufacturers
- Privacy Advocates
Sources
[1]Machine Learning PlusApplication DevelopersWhat Is the OpenAI Chat Completions API?
Read on Machine Learning Plus →
[2]arXivAI Infrastructure EngineersAttention Is All You Need
Read on arXiv →
[3]Factlen Editorial TeamAI Infrastructure EngineersSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Artificial Intelligence
See all →AI Safety Governance
OpenAI Dismisses Three Safety Researchers Over Alleged Confidential Data Sharing
6 sources
Frontier Models
Google Releases Gemini 4 Argon Frontier AI Model With 1-Million-Token Output Limit
7 sources
Frontier Models
Anthropic, OpenAI, and Xiaomi Launch Frontier Models as Claude Opus 5.5 Sets New Benchmark
5 sources
AI Architecture
The Core Mechanics of LLM Hallucination: Why Models Invent Facts and the Evidence on Mitigation
7 sources
Comments
Every angle. Every day.
Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.




