Skip to main content
ExplainerAI HardwareNvidia H100· 6 min read· in Artificial Intelligence

Single-Token Autoregressive Decoding Operates at One FLOP per Byte, Trapping Accelerator Tensor Cores Far Below the Roofline Ridge Point

During single-token generation, large language models load massive weight matrices from memory only to perform a single calculation per parameter. This structural asymmetry forces state-of-the-art AI accelerators to operate at a fraction of a percent of their theoretical compute capacity, fundamentally bottlenecking inference speed.

By Mateo Ramos

In short

  • Autoregressive decoding requires loading the entire model from memory for every generated token, resulting in an arithmetic intensity of just one FLOP per byte.
  • This 1:1 ratio falls 295 times below the Nvidia H100's compute-to-memory ridge point, leaving Tensor Cores idle for over 99 percent of the generation cycle.
  • Hardware upgrades like the H200 improve generation speeds not by adding compute power, but by increasing memory bandwidth to 4.8 terabytes per second.

The binding constraint for any processor to reach its advertised speed is that data must arrive at the compute units at least as fast as those cores can multiply it. For the current generation of large language models generating text one word at a time, this condition catastrophically fails.[3]

The architecture of autoregressive decoding forces the system to read the entire model from memory for every single step forward. Because each new word depends on the sequence that came before it, the generation process cannot be parallelized across time.[3]

The result is a structural arithmetic bottleneck that leaves the world's most advanced artificial intelligence accelerators sitting almost entirely idle. When an Nvidia H100 GPU generates a single token, its specialized Tensor Cores spend more than 99 percent of their cycles waiting for data to arrive.[3]

This phenomenon is known as the memory wall. It dictates the economics of artificial intelligence deployment, governs the design of next-generation hardware, and explains why generating a response is fundamentally slower than reading a prompt.[3]

The arithmetic of a single token

To understand the bottleneck, one must trace the exact mathematical operations required to generate a single token. A language model is essentially a massive collection of weight matrices.[3]

In a standard 7-billion-parameter model stored in 16-bit floating-point precision, those weights occupy roughly 14 gigabytes of space. During inference, the GPU must load every single one of those parameters from its High Bandwidth Memory into its on-chip static RAM.[2][3]

A 7-billion-parameter model requires moving 14 gigabytes of data for every single token generated.

For each parameter loaded, the processor performs exactly two mathematical operations: one multiplication and one addition. This ratio defines the workload's arithmetic intensity.[1][3]

Because the system loads two bytes of data to perform two floating-point operations, the arithmetic intensity of single-token decoding is exactly one FLOP per byte. This 1:1 ratio is the inescapable mathematical reality of batch-size-1 autoregressive generation.[3]

The roofline model and the ridge point

Hardware architects visualize this relationship using the roofline model, a framework that plots a processor's performance limits based on arithmetic intensity. The model forms a literal roof shape with a sloped ceiling and a flat top.[1]

The sloped section represents workloads limited by memory bandwidth, where performance scales linearly with data delivery. The flat horizontal line represents the absolute physical limit of the compute cores, where the processor is performing math as fast as its silicon allows.[1]

The intersection of these two lines is called the ridge point. It represents the exact arithmetic intensity required to keep the compute cores fully fed with data and escape the memory bottleneck.[1]

On an Nvidia H100 SXM, the theoretical peak compute for 16-bit operations is 989 teraflops. Its memory subsystem can deliver data at a maximum rate of 3.35 terabytes per second.[3]

The roofline model demonstrates the exact point where a processor transitions from being memory-bound to compute-bound.

Dividing the compute capacity by the memory bandwidth yields a ridge point of 295 FLOPs per byte. To fully utilize an H100, an algorithm must perform 295 operations for every byte it pulls from main memory.[3]

The utilization collapse

Single-token decoding operates at just one FLOP per byte, placing it 295 times below the H100's ridge point. The workload is pinned to the extreme far left of the roofline plot, deep in the memory-bound regime.[1][3]

The physical consequence of this ratio is a massive utilization collapse. The GPU loads weights at 3.35 terabytes per second, but because it does almost nothing with each byte, the compute units finish their work instantly.[3]

At a batch size of one, the peak compute utilization of an H100 is less than 0.4 percent. The processor is effectively acting as a highly expensive memory controller, with its 989-teraflop Tensor Cores starved of work.[3]

This explains why a 14-gigabyte model takes roughly seven milliseconds just to load its weights across a two-terabyte-per-second memory bus, while the actual computation requires only a fraction of a millisecond. The remaining time is pure latency.[3]

The prefill asymmetry

This memory-bound reality only applies to the decoding phase. When a user first submits a prompt, the model processes all the input tokens simultaneously in a phase called the prefill.[3]

Processing a prompt fully utilizes the GPU's compute capacity, while generating a response leaves the processor almost entirely idle.

During prefill, the GPU still loads the 14 gigabytes of weights once. However, it multiplies those same weights against every token in the prompt at the same time.[3]

If a user submits a 2,048-token prompt, the arithmetic intensity jumps to roughly 2,048 FLOPs per byte. This pushes the workload far past the 295-FLOP ridge point and onto the flat roof of the performance chart.[1][3]

In this compute-bound regime, the Tensor Cores run at maximum capacity, and memory bandwidth is no longer the limiting factor. This structural asymmetry explains why processing a 1,000-word prompt takes a fraction of a second, while generating a 1,000-word response takes several seconds.[3]

How batch size shifts the bottleneck

The primary lever engineers use to increase arithmetic intensity during generation is batching. By processing multiple user requests concurrently, the server can load the model weights once and apply them to several different output tokens simultaneously.[3]

Increasing the batch size from one to 32 raises the arithmetic intensity to 32 FLOPs per byte. While this is still well below the H100's ridge point, it effectively multiplies the throughput of the system by a factor of 32 without requiring additional memory reads.[3]

However, batching introduces its own constraints. Every concurrent sequence requires its own Key-Value cache to store the context of previously generated tokens.[3]

Recent hardware upgrades have focused entirely on increasing memory bandwidth to accelerate generation speeds.

As the batch size and context length grow, the cache consumes an increasing share of the GPU's memory capacity and bandwidth. Researchers have found that at large batch sizes, DRAM bandwidth saturation remains the primary bottleneck, stalling over half of the attention kernel cycles.[3]

Hardware and software interventions

Because raw compute power cannot solve a memory bandwidth problem, hardware manufacturers have shifted their focus to memory speed. The Nvidia H200 maintains the exact same 989-teraflop compute capacity as the H100, but upgrades the memory to the newer HBM3e standard.[2][3]

This new memory standard delivers 4.8 terabytes per second of bandwidth, a 43 percent increase over the H100. For memory-bound decoding workloads, this translates directly into a 43 percent increase in tokens generated per second.[2][3]

On the software side, quantization offers a mathematical workaround. By compressing 16-bit weights into 8-bit or 4-bit integers, engineers halve or quarter the amount of data that must cross the memory bus.[3]

An 8-bit model requires only half the memory bandwidth to load, effectively doubling the generation speed. While this introduces minor precision losses, it is currently the most effective software intervention for escaping the memory wall.[3]

Alternative architectures, such as speculative decoding, attempt to bypass the bottleneck entirely. These systems use a tiny, fast draft model to guess several tokens ahead, then use the large model's idle compute capacity to verify the entire batch in a single parallel step.[3]

Quantization halves the amount of data that must cross the memory bus, effectively doubling generation speed.

The fundamental physics of data movement dictate that memory bandwidth improves much more slowly than compute density. Over the past decade, artificial intelligence accelerators have seen exponential leaps in teraflops, while memory speeds have scaled linearly.[1][3]

Until a paradigm shift in hardware architecture allows memory and compute to physically merge, the single-token decoding process will remain trapped below the ridge point. The industry's most powerful processors will continue to spend their time waiting for data to arrive.[3]

How we did this

Method
Recomputation of hardware utilization by comparing theoretical peak matrix-multiplication throughput against memory bandwidth limits during batch-size-1 autoregressive decoding.
What we found
At batch size 1, an H100 spends over 99.6% of its compute cycles idle waiting for memory, operating at an effective utilization of less than 0.4% of its theoretical peak.
What we worked from
Limits of this analysis
This calculation assumes a pure batch-size-1 workload without speculative decoding, KV cache constraints, or network latency, which would further alter real-world utilization.

Key terms

Arithmetic Intensity
The ratio of mathematical operations performed for every byte of data loaded from memory.
Roofline Model
A visual graph used by hardware engineers to determine whether a program is limited by memory speed or compute power.
Ridge Point
The specific arithmetic intensity at which a processor's memory bandwidth perfectly matches its maximum compute capacity.
Tensor Cores
Specialized processing units inside modern GPUs designed specifically to multiply large matrices of numbers together rapidly.
High Bandwidth Memory
A 3D-stacked memory architecture that places RAM physically closer to the processor to achieve terabyte-per-second data transfer speeds.

Frequently asked

Why doesn't adding more GPUs make single-token generation faster?

Splitting a model across multiple GPUs reduces the memory load per chip, but introduces network communication delays. The time spent transferring data between the GPUs often consumes the latency gained by the smaller memory reads.

How does Apple's Unified Memory architecture handle this bottleneck?

Apple Silicon chips share memory between the CPU and GPU, offering up to 400 gigabytes per second of bandwidth. While highly efficient for consumer devices, this is still roughly eight times slower than an H100's dedicated HBM3, making generation proportionally slower.

Do diffusion models for image generation face the same memory wall?

No. Diffusion models perform dozens of complex denoising steps on the same loaded weights, pushing their arithmetic intensity much higher. They are typically compute-bound, which is why they benefit directly from increased teraflops.

Viewpoints in depth

Hardware Architects

Engineers focused on increasing physical memory bandwidth and developing new interconnect standards.

Silicon engineers view the memory wall as a physical packaging problem. Because High Bandwidth Memory must be placed extremely close to the compute die to achieve terabyte-per-second speeds, thermal limits and physical space constrain how much bandwidth can be added. Their primary roadmap involves 3D stacking and silicon photonics to widen the data highway without melting the processor.

Algorithm Researchers

Researchers focused on bypassing the autoregressive bottleneck through software innovation.

Software researchers argue that waiting for hardware to solve the memory wall is a losing battle, given that compute scales faster than bandwidth. Their focus is on altering the decoding mechanism itself. By developing speculative decoding and non-autoregressive generation techniques, they aim to artificially inflate the arithmetic intensity of inference so that the existing idle compute cycles can be put to use.

Edge Deployment Advocates

Engineers prioritizing extreme quantization to run models on bandwidth-constrained consumer hardware.

Engineers deploying models to phones and laptops face memory bandwidths measured in gigabytes, not terabytes. This camp prioritizes extreme quantization, pushing models down to 4-bit or even 2-bit precision. They accept the slight degradation in reasoning quality as a necessary trade-off to fit the workload within the strict memory bandwidth limits of consumer-grade silicon.

Hardware Architects 35%Algorithm Researchers 35%Edge Deployment Advocates 30%
Hardware Architects
Engineers focused on increasing physical memory bandwidth and developing new interconnect standards.
Algorithm Researchers
Researchers focused on bypassing the autoregressive bottleneck through software innovation.
Edge Deployment Advocates
Engineers prioritizing extreme quantization to run models on bandwidth-constrained consumer hardware.

Perspectives this story doesn't cover

  • Cloud Infrastructure Providers
  • Energy Grid Operators

Sources

Source coverage

3 outlets

3 viewpoints surfaced

Hardware Architects 35%Algorithm Researchers 35%Edge Deployment Advocates 30%
  1. [1]WikipediaHardware Architects

    Roofline model

    Read on Wikipedia →
  2. [2]WikipediaHardware Architects

    High Bandwidth Memory

    Read on Wikipedia →
  3. [3]Factlen Editorial TeamAlgorithm Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.