When is LLM inference actually memory-bound?
TL;DR. A GPU has two speeds: how fast it can do math, and how fast it can pull numbers out of its own memory. To keep its math units fed, a B200 needs roughly 280 operations of useful work for every byte it fetches (other recent GPUs land in the same 150–300 regime); generating text one token at a time manages about one. That mismatch is what people mean when they say inference is "memory-bound", and it's real, but it is a property of an operating point, not of LLMs: batching changes it, long conversations change it back, and a measurement far below both of the hardware's limits means the simple model has stopped describing your system at all. You can tell which situation you're in with arithmetic that fits on an index card. This post builds that card from zero.
1. A number that bothered me
While benchmarking Shard I measured how fast an 8-billion-parameter Llama model generates text on a B200, a high-end Blackwell datacenter GPU, using the standard HuggingFace library. The answer: about 29 tokens per second. (A token is a word-sized chunk of text, so that is roughly 20 words per second.)
Is that good? The honest answer is that most people, including people who work with these models daily, have no idea how to check. The spec sheet advertises quadrillions of operations per second; generation ran at 29 tokens per second. How can both numbers describe the same machine? Somewhere between them, intuition runs out.
This post is the missing check. By the end, "29 tokens/sec" will read as what it is: about 6% of a ceiling you can compute on the back of an envelope, and the reasons why come down to counting two things.
2. A GPU has two speeds
Everything below follows from one piece of hardware reality. A GPU has:
- Math speed: peak arithmetic throughput of its tensor cores, counting a multiply-add as two operations. At dense BF16/FP16 precision, the precision used for the comparisons in this post, a B200 does about 2,250 trillion operations per second.
- Memory speed (bandwidth): how many bytes it can move from its onboard memory to its math units per second. A B200 moves about 8 trillion bytes (8 TB) per second.
Both numbers sound enormous until you divide them: ~280 math operations per byte fetched. If a task doesn't do a couple hundred operations with every byte it fetches, the math units sit idle, waiting on memory. Think of a kitchen that can plate 280 dishes in the time the pantry delivers one basket of ingredients: unless each basket feeds hundreds of dishes, the cooks are just standing around.
The ratio of work-per-byte a task achieves decides everything, and it has a name, arithmetic intensity. Comparing it against the hardware's ratio is called the roofline model[roofline]:
$$ \text{time per step} \;=\; \max\!\left( \frac{\text{math to do}}{\text{math speed}},\;\; \frac{\text{bytes to fetch}}{\text{memory speed}} \right) $$Whichever term is bigger is your wall. Here's the balance point for the GPUs people actually run (using honest "dense" math figures; datasheets often quote 2× these via a sparsity feature LLMs don't use):
| GPU | math speed | memory speed | ops per byte at balance |
|---|---|---|---|
| A100 (2020) | 312 T/s | 2.0 TB/s | ~150 |
| H100 (2022) | 989 T/s | 3.35 TB/s | ~300 |
| H200 (2024) | 989 T/s | 4.8 TB/s | ~200 |
| B200 (2025) | ~2250 T/s | 8.0 TB/s | ~280 |
| RTX 4090 | 165 T/s | 1.0 TB/s | ~160 |
Despite huge absolute gains on both axes, five years of hardware keeps the balance point in the same regime, 150–300 operations per byte: math and memory have been scaling together. So the question for any workload, on any of these chips, is the same: how many operations do you do per byte you fetch?
3. Generating one token: read everything, use it once
A language model is, physically, a very large pile of numbers, its weights. Llama-3.1-8B is 8 billion weights; stored at 2 bytes each, that's 16 GB. Generating the next token means taking the current token's hidden state, a vector of a few thousand numbers, and passing it through essentially all 8 billion weights, while attention reads the cached state of the earlier tokens (hold that thought for §6). For a dense model like this one, every weight participates in every token; there is no shortcut where the model only consults the relevant part. (Mixture-of-experts models change exactly this accounting, activating only a subset of their weights per token, which is a large part of why they exist.)
Now count, per generated token:
- Bytes fetched: all the weights, once ≈ 16 GB.
- Math done: each weight is used in roughly two operations (one multiply, one add) ≈ 16 billion operations.
That's one operation per byte, on hardware balanced for ~280, which means that even with perfect use of the memory system, this shape of work can touch at most ~0.4% of the chip's peak math before bandwidth runs out. No clever kernel changes that for conventional single-stream generation; the shape of the task, a small vector against a giant matrix, simply doesn't contain enough math per byte. (Changing the shape is a different story, §5.) So the binding ceiling is the memory term. It's an idealized ceiling, real decoding also moves attention state and activations, and no kernel sustains 100% of peak bandwidth, but it's the right yardstick:
$$ \frac{8\,\text{TB/s memory speed}}{16\,\text{GB of weights per token}} \approx 500 \text{ tokens/s} $$against a math ceiling of 2,250T ÷ 16B ≈ 140,000 tokens/s. Single-stream generation is a memory benchmark wearing a chatbot costume.
Two useful consequences fall out immediately:
- For low-batch decode, bandwidth matters more than FLOPs. For a bandwidth-bound workload at comparable efficiency, the H200's ceiling is ~43% higher than the H100's at identical math speed, because 4.8/3.35 ≈ 1.43.
- Shrinking the weights raises the ceiling. Quantization, storing each weight in 1 byte or half a byte instead of 2, is usually sold as a way to fit models in smaller GPUs. But look at the equation: halve the weight bytes and you double the weight-bandwidth ceiling. That's why quantization can be a speed optimization, not just a capacity one.
4. Reading your prompt is the opposite problem
Before the model generates anything it has to process your prompt, and here the arithmetic flips. All the prompt's tokens can go through the weights together: fetch a weight once, use it against 2,000 tokens. For the big matrix multiplies that dominate the work at ordinary prompt lengths, that's thousands of operations per byte, way past the ~280 balance point, so prompt processing is normally math-bound; the crossover sits at a prompt of a few hundred tokens. (At very long contexts the attention step, whose work grows with the square of the length, starts to complicate the picture.)
So every request has two phases with opposite bottlenecks:
This is why "time to first token" and "time per token after that" respond to completely different fixes, and why a provider's long-prompt pricing and generation pricing can differ: they're spending different resources on each.
5. Batching makes the weight reads dramatically cheaper per user
Here's the observation the entire economics of AI serving rests on. The 16 GB of weights fetched to generate your next token could just as easily be used to generate everyone's next token: the fetch is the expensive part, the math units are idle anyway. Serve 100 conversations at once and you do 100× the math for the same memory traffic, 100 ops per byte instead of 1.
Ignore the KV cache for one paragraph, weights only. In that model, total output scales almost linearly with batch until compute pushes back near the ~280 balance point, a couple hundred simultaneous streams per GPU, at modest cost to each user's speed. (Not free: bigger batches mean more math per step, and real systems trade a little latency to form batches at all. But each marginal user rides on bandwidth that was already being spent.) This amortization is one big reason why tokens from a shared API cost fractions of a cent while a dedicated GPU costs dollars per hour, and why "we'll give every user their own replica" business plans die on contact with an invoice.
So batch up to ~280 and the GPU is perfectly balanced? Not quite. Something grows alongside the batch and eats the win.
6. Each conversation drags its memory around
To generate the next token, the model doesn't just need its weights; it needs its working notes on every token so far in this conversation, what each earlier word was, in a form the attention mechanism can look back at. Those notes are the KV cache, and for Llama-3.1-8B they cost 128 KB per token of conversation. An 8,192-token conversation carries exactly 1 GiB of notes (1.07 GB).
The weights are shared by everyone in the batch. The notes are not, every conversation re-reads its own, for every token generated. So the bytes-per-token bill looks like:
| fetched per generated token | bytes | shared across the batch? |
|---|---|---|
| weights | 16 GB | yes |
| one conversation's notes @ 8K history | 1.07 GB | no |
At what batch size do the notes outweigh the model? $16\,\text{GB} \div 1.07\,\text{GB} \approx 15$. Fifteen long conversations, and the GPU now spends more of its memory budget re-reading histories than reading the model, and this part doesn't get cheaper with scale, because every added conversation adds its own notes. The amortization lever from §5 stalls long before ~280. And the notes keep growing: by thirty 8K conversations, the cache alone is already twice the size of the model weights.
Once you see this, the last three years of model-architecture news reads differently. GQA[gqa] (which Llama already uses) cut the notes 4× versus the full-attention equivalent. DeepSeek-V2's MLA[mla] reported a 93.3% KV-cache reduction against its own baseline. Quantized caches, cache eviction, and compression schemes (Shard is our 10× attempt) all exist because the KV cache is the one memory cost that serving at scale makes worse.
7. The third situation: neither wall
Back to the number that started this: 29 tokens/sec, against a memory ceiling of ~500 and a math ceiling of ~140,000. It's nowhere near either wall, and strictly speaking, that means the two-resource model has stopped describing the system. Many things can live in that gap: kernel-launch latency, matrix shapes too small to fill the chip, synchronization, scheduling, cache behavior. You have to profile to know. When we profiled this baseline (the Shard write-up has the traces), it was the classic case: tens of thousands of tiny kernel launches and tensor copies per generation, dispatched one by one from a Python loop, each with a fixed cost the GPU waits through. Horace He's name for this regime fits: overhead-bound[brrr].
The deceptive part: monitoring tools can report "GPU utilization: 100%" in this state, because that metric measures the fraction of time during which at least one kernel was executing, not how much of the silicon it used[util]. A GPU can report 100% utilization while achieving a few percent of its peak arithmetic throughput.
The fixes are software: fused operations, captured execution graphs (CUDA Graphs) replayed in place of thousands of individual launches, and purpose-built serving engines (vLLM[vllm], TensorRT-LLM, SGLang) that keep per-token dispatch cost off the critical path. Optimized runtimes close a large fraction of the gap to the memory ceiling, and however much of it they close, the point stands: the distance from 6% is the same GPU, actually asked to work.
8. The index card
Everything above compresses to a procedure:
- Compute two ceilings (aggregate tokens/sec, batch size $B$). Memory: $B$ × bandwidth ÷ model bytes, since the weight read is shared, this one grows with batch. Math, to first order: math speed ÷ (2 × parameter count), and this one doesn't grow with batch: batching spends idle math, it doesn't mint more. (At very long contexts, add the attention work, which grows with context, to the math bill too.) Your ceiling is the smaller of the two. Thirty seconds with a spec sheet; use dense figures. Completed, the memory ceiling is $$ \text{memory ceiling} \;\approx\; \frac{B \cdot \text{bandwidth}}{W + B \cdot K \cdot L} \;\text{tokens/s} $$ where $W$ is weight bytes, $K$ is KV bytes per token per stream (128 KB here), and $L$ is context length. Small $B$: the weights dominate and batching helps almost linearly. Large $B$ or $L$: the $BKL$ term takes over and batching stops buying much. That one line is §§3, 5 and 6.
- Measure prompt processing and generation separately; they're different workloads.
- Compare. Generation near the memory ceiling: genuinely memory-bound; your levers are quantization, batching, speculative decoding[spec], or a higher-bandwidth card. Prompt phase near the math ceiling: buy FLOPs or shorten prompts. Far below both: the simple roofline model isn't describing your system. Profile launches, kernel efficiency, and synchronization before you shop; more peak FLOPs or bandwidth alone may not move it.
- Long conversations or big batches? Watch the $BKL$ term; the wall quietly moves from the weights to the notes, and the fix list changes with it.
"Inference is memory-bound" is true the way "planes fly west slowly" is true: a real effect, specific conditions, and knowing which conditions is the whole skill. The arithmetic fits on an index card. It's the difference between optimizing the thing that's slow and optimizing the thing that's famous for being slow.
Numbers: Llama-3.1-8B-Instruct (8.03B weights, fp16), NVIDIA dense-BF16 datasheet figures. The 29 tok/s measurement is the FP16 baseline from the Shard benchmark: HuggingFace transformers, eager mode, batch 1, B200. KV cache: 2 · 32 layers · 8 KV heads · 128 dims · 2 bytes = 128 KB per token.
References
- Roofline: Williams, Waterman, Patterson. Roofline: An Insightful Visual Performance Model for Multicore Architectures. CACM 2009.
- Compute / memory / overhead trichotomy: He. Making Deep Learning Go Brrrr From First Principles. 2022. horace.io/brrr_intro.
- Inference arithmetic, in more depth: kipply. Transformer Inference Arithmetic. 2022. kipp.ly; Pope et al. Efficiently Scaling Transformer Inference. arXiv:2211.05102.
- GQA: Ainslie et al. arXiv:2305.13245.
- MLA: DeepSeek-AI. DeepSeek-V2. arXiv:2405.04434.
- Speculative decoding: Leviathan, Kalman, Matias. arXiv:2211.17192; Chen et al. arXiv:2302.01318. Uses the idle math units to check several drafted tokens in one weight-fetch; the memory-boundness of generation is exactly why it works.
- vLLM: Kwon et al. Efficient Memory Management for LLM Serving with PagedAttention. SOSP 2023. arXiv:2309.06180.
- GPU "utilization": NVML defines it as the percentage of the sampling period during which one or more kernels were executing, not achieved throughput.