Section 01

What is an LLM?

And what does "inference" mean?

Sources: The Llama 3 Herd of Models — Grattafiori et al., 2024; Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023

When you type something into ChatGPT, Claude, or any modern chat assistant and watch words appear one by one, something concrete is happening on a computer. A program is reading your text, turning it into numbers, doing an enormous amount of arithmetic on those numbers, and turning the result back into more text. That program is a Large Language Model (LLM)LLMLarge Language Model — a neural network trained on huge text corpora to predict the next token given previous tokens.See in glossary → — a neural networkneural networkA function built by stacking many simple operations — mostly matrix multiplies with nonlinearities between them — whose behavior is shaped by tuning billions of internal numbers (its parameters) from data.See in glossary →, a mathematical function whose adjustable numbers are learned from examples — and the act of running it to produce output is called inferenceinferenceRunning a trained model to produce outputs. Training learns the weights once; inference uses them many times.See in glossary →.

Running that model efficiently is a substantial engineering task. An inference engine such as vLLMvLLMAn open-source LLM inference engine, originally from UC Berkeley, that introduced paged attention and is now one of the most widely used serving systems for open-weight models.See in glossary →, software that runs language models and manages incoming requests, has to keep an expensive graphics processing unit (GPU)GPUGraphics Processing Unit — a processor designed to perform many calculations in parallel, widely used for training and running neural networks.See in glossary → busy while many people use it at once. A GPU is a chip designed to perform many arithmetic operations in parallel.

Training vs inference

A neural network is, at heart, a giant function. It takes numbers in, multiplies them by other numbers (called parametersparametersThe numbers (weights) inside a model that get adjusted during training. A “7B model” has 7 billion of them.See in glossary →, or weights), passes the results through some simple nonlinear functionsnonlinear functionA function whose output isn't just a scaled, shifted copy of its input — e.g. ReLU, GELU, sigmoid. Stacking nonlinearities between matrix multiplies is what lets a neural net represent anything more interesting than scaling and rotation.See in glossary →, which do more than multiply inputs by fixed weights and add them, and produces numbers out. The interesting trick is that the parameters are learned from data. We start with random parameters, show the network billions of examples (“here is some text: predict what comes next”), and slowly nudge the parameters so it does better. That nudging process is called training.

Training a particular model version is expensive. Frontier models cost tens of millions of dollars and run for months on tens of thousands of GPUs. Once that version is deployed, its parameters are frozen (they are just a big file of numbers, hundreds of gigabytes for a flagship model) and serving uses those fixed parameters to answer users’ prompts. Developers can later train or fine-tune a new version, but that happens outside live inference.

What does an LLM actually compute?

The language models considered here generate text by repeatedly predicting its next piece. That piece is a tokentokenThe atomic unit of text the model sees. Roughly a word-fragment — “tokenization” is a piece of text → list of token IDs.See in glossary →, often a word or a word fragment. Their training task is given preceding text, predict the next token. The model is shown a chunk of text and asked: what comes next?

Once you have a really good next-token predictor, you can string it together: predict the next token, append it to the input, predict the next token after that, and so on. That loop (generate one token, feed it back in, generate another) is what we call autoregressiveautoregressiveGenerating one token at a time, where each new token is conditioned on every token that came before it.See in glossary → generation. This is how most text chat assistants generate their responses.

The piece you type in is called the promptpromptThe input text fed to the model — what you want it to continue or respond to.See in glossary →. The text the model generates in response is the completioncompletionThe text the model generates in response to a prompt.See in glossary →. Everything you see being typed out token by token in a chat UI is the model running through its autoregressive loop. Each generated token requires another complete forward passforward passRunning inputs through the network to produce outputs (logits) and the loss, caching intermediate activations that backpropagation will need.See in glossary →, the calculation that carries inputs through the model to its outputs: on the order of (10^11) arithmetic operations for a densedense modelA model that uses all its main parameter blocks for each input, rather than selecting only some expert blocks.See in glossary → 70B-parameter model, one that uses all its main weights for each token, before counting attention work that grows with the context.

Why is this hard?

Two constraints make this repeated prediction expensive:

  1. The model is gigantic. A flagship open-weights model like Llama-3-70B has 70 billion parameters. At 16-bit precision, meaning each parameter occupies 16 binary digits, or 2 bytes, that is about 140 gigabytes (GB) of weights. That does not fit on an 80 GB H100, and even a 141 GB H200 leaves essentially no space for the KV cacheKV cacheThe stored keys and values from all past tokens, so attention at step t only needs to compute Q for the new token.See in glossary →, stored intermediate results that let the model reuse earlier calculations, and other temporary data. Even the smaller 8B model is 16 GB. Every single token you generate requires reading every parameter of a dense model at least once from GPU memory. During token-by-token generation with small batchesbatchThe group of training examples used for one gradient estimate. Bigger batches reduce gradient noise but use more memory and compute per step.See in glossary →, groups of inputs processed together, the chip is usually limited by how quickly it can move those weights from memory, rather than by its arithmetic throughput. That asymmetry shapes many decisions in a serving system.

  2. Many users want answers at the same time. If you decode one request one token at a time, GPUs are usually underutilized: the step reads the model’s weights but performs only one token’s worth of work with them. Real systems pack many requests together so that one read of the weights serves many users, while juggling the fact that those requests have wildly different prompt lengths, completion lengths, and arrival times.

A modern inference engine like vLLM is, fundamentally, an answer to both problems. Managing memory is central to its job.

From text to a running service

The problem has four main parts:

  1. Foundations (sections 2–10). How text becomes numbers, how the model combines information from different positions, and how its calculations produce a prediction.

  2. How inference actually runs (sections 11–14). Processing the prompt, generating the response, storing reusable results, and grouping requests so they share work.

  3. vLLM internals (sections 15–18). Techniques for organizing stored results, reusing work across requests, and predicting several tokens at once. These affect throughputthroughputTotal tokens generated per second across all concurrent requests. Often traded against per-request latency.See in glossary →, the amount of work completed per second, and latencylatencyThe elapsed time between a request or operation starting and the relevant result becoming available.See in glossary →, the time a user waits for a result. Their benefit and achievable concurrency depend on the model, hardware, request mix, and configuration.

  4. Scaling out (sections 19–21). What changes when one GPU isn’t enough.

Every request begins with the same operation: turning the input text into numbers.