How Large Language Models Actually Work
From tokens to attention to next-word prediction: a ground-up tour of the machinery behind modern LLMs.
Debanjan Saha
· 4 min read
On this page
A large language model does one thing: given some text, it predicts what comes next. Everything impressive about it, from writing code to summarising contracts, falls out of doing that one thing extremely well, at enormous scale. This article walks through the pieces that make it possible.
Text becomes tokens
Models don’t read characters or words. They read tokens: chunks of text drawn from a fixed vocabulary, typically tens of thousands of entries. A common scheme is byte-pair encoding, which starts from single bytes and repeatedly merges the most frequent adjacent pairs until the vocabulary is full.
The result is that common words are a single token, rare words split into several, and numbers or unusual names can fragment awkwardly. That fragmentation explains a surprising number of model quirks, such as trouble counting the letters in a word.
Tokens become vectors
Each token ID is looked up in an embedding table, producing a vector of a few thousand numbers. Nothing about the vector is hand-designed. Training nudges it so that tokens used in similar ways end up near each other in this high-dimensional space.
Because the architecture processes all positions in parallel, it must also be told where each token sits. Positional information is injected into those vectors, using schemes such as rotary position embeddings.
Attention: letting tokens look at each other
The core of the transformer, introduced in the 2017 paper “Attention Is All You Need”, is the attention mechanism. For every token, the model computes three vectors: a query (what am I looking for?), a key (what do I offer?) and a value (what do I pass along if selected?).
Each token’s query is compared with every earlier token’s key. The scores are normalised into weights, and the output is a weighted average of the values:
import numpy as np
def attention(Q, K, V):
d_k = Q.shape[-1]
scores = Q @ K.T / np.sqrt(d_k) # how relevant is each key to each query?
mask = np.triu(np.ones_like(scores), k=1) # causal: no peeking at future tokens
scores = np.where(mask == 1, -np.inf, scores)
weights = np.exp(scores - scores.max(-1, keepdims=True))
weights /= weights.sum(-1, keepdims=True) # softmax
return weights @ V
In a real model this runs many times in parallel as multi-head attention, so different heads can specialise in different relationships, such as syntax, coreference or long-range structure.
Layers upon layers
A transformer block pairs attention with a small feed-forward network, wrapped in residual connections and normalisation. Dozens of these blocks are stacked. Early layers tend to capture local patterns; later ones build up more abstract representations. The final layer’s output is projected back onto the vocabulary to give a score for every possible next token.
Predicting the next token
Those scores are converted to probabilities, and one token is chosen. How it is chosen matters:
- Greedy decoding always takes the most probable token. It is deterministic but can be dull and repetitive.
- Sampling draws from the distribution. A temperature parameter sharpens or flattens it; higher values give more varied output.
- Top-k and top-p (nucleus) sampling restrict the draw to the most plausible candidates.
The chosen token is appended to the input and the process repeats, one token at a time. That is all “generation” is.
Training: where the knowledge comes from
Training happens in stages.
- Pretraining. The model reads a vast quantity of text and is trained to minimise the error of its next-token predictions. This is self-supervised: the text supplies its own labels. It is by far the most expensive stage.
- Supervised fine-tuning. The base model, which merely continues text, is trained on curated examples of instructions and good responses so it behaves like an assistant.
- Preference tuning. Methods such as reinforcement learning from human feedback, or direct preference optimisation, push the model toward answers people rate as better.
What models can and can’t do
Understanding the mechanism clarifies the failure modes:
- Hallucination is not a bug in a lookup. The model produces plausible continuations, and plausible is not the same as true.
- The context window is the model’s only working memory. Anything outside it is invisible unless supplied again.
- Knowledge is frozen at the end of training, which is why retrieval and tools matter.
None of this makes LLMs less useful. It makes them predictable, and predictability is what lets you build reliable systems around them.
Further reading
Start with the original transformer paper, then Andrej Karpathy’s “Let’s build GPT” video, which implements the whole thing in a few hundred lines.
Discussion
Comments (Giscus) will appear here. Set
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORYandPUBLIC_GISCUS_CATEGORY_IDto enable them.