How large language models work.
published 2026-08-30 · revised 2026-09-29
A language model is a calculation that reads the words written so far and gives every word it knows a score for how likely that word is to come next. One word is then picked according to those scores, added to the text, and the calculation runs again. This page starts from a sentence with its last word missing and builds the calculation up in steps from there. By the end you will be able to say what a large language model computes when it answers you, and why its answer is picked rather than derived.
contents /
the missing word /
Here is a sentence with its last word missing. Read it and notice which word comes to mind.
It was the best of times, it was the worst of ____
The word is “times”. If you know the book, you remembered it. If you do not, “times” came to mind anyway, because the sentence had already used the word once; in that case, nothing but the words on the page suggested it. Either way, a few words came forward, one of them ahead of the rest, and you took that one.
Claude Shannon measured how well people do this, in a 1951 paper called “Prediction and Entropy of Printed English”. He gave a subject a passage they had not read and asked them to guess it one letter at a time. After each guess the subject was told whether it was right and, if it was wrong, what the letter was. The correct text so far stayed in front of the subject throughout, so every guess was made knowing everything that came before it. In the run he printed, 89 of 129 letters were guessed correctly, and the errors fell where you would expect: they “occur most frequently at the beginning of words and syllables where the line of thought has more possibility of branching out”.
The experiment measures how much of a text a person can supply from what came before it. Shannon opens that section of the paper by naming the ability being measured: familiarity with a language’s words, idioms, clichés and grammar is what lets somebody fill in missing or incorrect letters in proof-reading, or “complete an unfinished phrase in conversation”. The sentence at the top of this page is an unfinished phrase of that kind, and you completed it.
The boxed equations on this page are the actual definitions. Each box is followed by a line in plain English that says the same thing, so you can skip the boxes and lose only the notation. In most of the figures you can pick a word or press a button, and the numbers change as you do.
The rest of this page finishes that sentence with arithmetic, one method at a time. Each method improves on the one before it and still fails somewhere, and the next method starts from that failure.
counting what follows /
The first method, and the simplest, is to count. Go through a body of text and record, for every word in it, which words came next and how often. Finishing a sentence is then a lookup: find the sentence’s last word in the table, and answer with the follower that was counted most often.1
The body of text here is one paragraph, and you have already finished its opening words.
It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness, it was the epoch of belief, it was the epoch of incredulity, it was the season of Light, it was the season of Darkness, it was the spring of hope, it was the winter of despair, we had everything before us, we had nothing before us, we were all going direct to Heaven, we were all going direct the other way—in short, the period was so far like the present period, that some of its noisiest authorities insisted on its being received, for good or for evil, in the superlative degree of comparison only.
— Charles Dickens, “A Tale of Two Cities”, 1859 · 119 words, 58 of them different
Figure 1 shows what followed each word in that paragraph. Pick a word to see its list. Then press “take the commonest”, and the figure writes a sentence: it takes the commonest follower, looks that word up in turn, and repeats.
- the
- of
- was
- it
- we
- times
- age
- epoch
- season
- had
- before
- us
- were
- all
- going
- direct
- in
- period
- its
- for
- best
- worst
- wisdom
- foolishness
- belief
- incredulity
- light
- darkness
- spring
- hope
- winter
- despair
- everything
- nothing
- to
- heaven
- other
- way
- short
- so
- far
- like
- present
- that
- some
- noisiest
- authorities
- insisted
- on
- being
- received
- good
- or
- evil
- superlative
- degree
- comparison
- only
after the · 14 times in the paragraph · 11 different words followed it
- age2
- epoch2
- season2
- best1
- worst1
- spring1
- winter1
- other1
- period1
- present1
- superlative1
- nothing else followed “the”
the
press to write a sentence, one word at a time
In mathematical terms the table holds one number for every pair of words:
read: the chance of a word, given the one before it, is how often the pair occurred over how often the first word occurred.
The table holds no information about what the words mean: only which word followed which, and how often.
The same idea appears in Shannon’s 1948 paper “A Mathematical Theory of Communication”, as the sixth in a series of approximations to English: “Second-order word approximation. The word transition probabilities are correct but no further structure is included.” Shannon made his sample by hand, from a book and the following rule.
Open the book at a random page and pick a word on it at random. Write that word down. Then open the book somewhere else, read forward until you meet the same word, and write down the word that comes after it. Open the book a third time and look for that new word, and so on. Every word you write depends on wherever the word before it happened to turn up. Shannon describes the rule for letters, and says the word samples were produced the same way.
Here is the start of Shannon’s sample, in his capitals:
THE HEAD AND IN FRONTAL ATTACK ON AN ENGLISH WRITER THAT THE CHARACTER OF THIS POINT IS THEREFORE ANOTHER METHOD FOR THE LETTERS
In short stretches the sample is nearly English, but as a whole it is about nothing. Shannon says how far the structure reaches: samples like this hold together “out to about twice the range that is taken into account in their construction”. The pairs this sample was built from are two words long, so twice that is about four words, and he reports that runs of four or more words “can easily be placed in sentences without unusual or strained constructions”. Some runs go further. He points to ten words of the sample, “attack on an English writer that the character of this”, and calls them “not at all unreasonable”.
Shannon did not present any of this as prediction. He never uses the phrase “language model” and never talks about guessing a next word; the samples are there to show how closely a made-up process can be made to imitate English.
The book method also differs from figure 1’s rule. Reading forward until you meet the word you last wrote down hands you each follower in proportion to how often it followed, so a word that followed twice as often turns up twice as often. In other words, Shannon draws a follower at random, where figure 1 always takes the commonest. The difference between the two is the subject of a later section, picking a word.
The counts from the Dickens paragraph behave much as Shannon described. Start at “it” and take the commonest follower every time, and the table writes it was the age of times. After that the same six words repeat for ever, because “times” is only ever followed by “it”, and “it” only ever by “was”.
Dickens wrote it was the age of wisdom. The first four words the table wrote are right and the fifth is wrong, and the error happens at “of”. The table remembers only that one word, and across the whole paragraph “of” is followed by “times” twice and by ten other words once each, so “times” wins. The word that would have decided the answer is “age”, but “age” sits two words back, and the table only ever looks at one.
There is also a smaller failure. Three words follow “the” twice each (“age”, “epoch” and “season”), so “take the commonest” does not name a single word, and figure 1 breaks the tie in favour of whichever appeared first. Break the tie another way and the same paragraph produces a different sentence.
“Times” is also the word you supplied at the top of this page, here turning up in the wrong clause. One word of memory is not enough, and the obvious fix is to remember more than one. The next section tries that and counts what it costs, starting with how much of the one-word table the paragraph already leaves empty.
when the table explodes /
Only 10 of the paragraph’s 58 words were ever followed by more than one word: 47 go exactly one place, and the last word goes nowhere.
That is why most of the lists in figure 1 have empty space below them. The paragraph uses 58 different words, so there are 58 × 58 = 3,364 possible pairs of them, and the paragraph contains 85. The other 3,279 pairs never occur, and a table of counts is normally about this empty.
The obvious fix, from the end of the last section, is to remember two words instead of one, and then three. Let’s try that on the sentence the table got wrong. With two words of memory the table looks up “age of” instead of “of”, and the paragraph contains “age of” twice: once before “wisdom” and once before “foolishness”. The table therefore stops answering “times”, but it does not start answering “wisdom” either, because the two followers are level at one occurrence each and only the tie-break rule separates them. The extra word of memory turned one wrong answer into a tie.
In mathematical terms, remembering only the last few words is an assumption, and everything later on this page inherits some version of it:
read: the next word really depends on the whole sentence so far; the count treats it as depending on the last few words only.
The next word depends on everything said so far, but no table could hold a count for everything that might have been said, so the table keeps the last few words and treats the rest as though they were not there. The only choice left is how many words to keep. The field names the result after the length of the sequence being counted, not the number of words remembered, and calls it an n-gram: one word of memory plus the next word makes a two-word sequence, a bigram, and two words of memory make three-word sequences, trigrams.
However, every extra word of memory multiplies the size of the table by the vocabulary (the number of different words in the text), while the paragraph stays the same length.
Figure 2 shows the result. In the grid, each of the 3,364 possible pairs has one small square: the row gives the first word, the column gives the second, and a square is inked only where that pair occurs in the paragraph. Those 85 marks are everything the table can learn from the paragraph. Underneath is the same count for longer sequences.
| memory | sequences possible | seen | one seen in |
|---|---|---|---|
| 1 word | 3,364 | 85 | 40 |
| 2 words | 195,112 | 96 | 2,032 |
| 3 words | 11,316,496 | 106 | 106,759 |
| 4 words | 656,356,768 | 110 | 5,966,880 |
Read that table a column at a time. Sequences possible counts every sequence the table would need an entry for. One word of memory means two-word sequences, so 58 × 58 = 3,364; each extra word of memory multiplies by 58 again, giving 195,112, then 11,316,496, then 656,356,768. Seen barely moves: 85, 96, 106, 110. The count rises slightly because longer sequences repeat less often, so more of them are distinct. The last column divides the first by the second: with one word of memory the paragraph contains one possible sequence in 40, and with four, one in 5,966,880.
A larger body of text has the same problem. Yoshua Bengio and three colleagues open their 2003 paper, “A Neural Probabilistic Language Model”, by stating the problem at full size, before proposing a way round it. Take a vocabulary of 100,000 words and try to count sequences ten words long: that is 100,000 multiplied by itself ten times, or about 10⁵⁰ sequences, each needing its own number. The paper calls those numbers free parameters: numbers nobody knows in advance, which have to be worked out from text. The paper names the difficulty the curse of dimensionality and states it in one sentence: “a word sequence on which the model will be tested is likely to be different from all the word sequences seen during training”.
The empty entries have a standard remedy, decades old. When a sequence has never been seen, fall back on a shorter one that has; and shave a little off every count you did see, so that something is left over for the ones you did not. This is called smoothing, and Jurafsky and Martin give it a section of their textbook. Smoothing makes the table usable, but it leaves a more serious fault in place.
The 2003 paper shows that fault with a pair of sentences. Suppose the text you counted over contained “The cat is walking in the bedroom”, and ask about “A dog was running in a room”, which it did not contain. The paper argues that having seen the first sentence should make the second almost as likely, “simply because ‘dog’ and ‘cat’ … have similar semantic and grammatical roles”. A table of counts cannot make that connection. The table treats the two as unrelated entries, the second with a count of zero, because nothing in a table of counts relates one word to another. A bigger table would have the same fault, so the fix has to replace the lookup with a calculation.
words as numbers /
The first step towards that calculation is to give every word a short list of numbers: two, say, or a hundred. The numbers give the word a position, the way a point on a map has one, and words used in similar ways can sit near each other. A list of numbers like this is called a vector, and a word’s vector is called its embedding. With vectors there is nothing to look up: a calculation reads the positions of the words it has been given and returns a score for every word that might come next.
The 2003 paper lists three steps: give each word a vector, work out the prediction from those vectors, and work out the vectors and the calculation at the same time, from the same text.
The third step is what captures the likeness that a table of counts could not see. Nobody writes down that “cat” and “dog” are alike. The two words turn up in similar company, so the numbers that make the calculation come out right end up placing them in similar positions, and once they are close, a sentence containing “cat” counts as evidence for the same sentence with “dog” in its place. The paper states this as a property of the method: a sequence never seen before still scores highly “if it is made of words that are similar … to words forming an already seen sentence”.
The calculation is short enough to write out. The vectors of the input words go in side by side, so two words of two numbers each make four numbers. In the middle is a row of units, and each unit does the same small job: it multiplies each of the four incoming numbers by a number of its own, adds the four results together, adds one more number of its own, and squashes the total into the range −1 to 1. A row of units like this is called a layer.
Each candidate word then multiplies and adds the layer’s numbers in the same way, with numbers of its own but without the squashing, and the result is one score per candidate word. The multipliers and the added numbers are what this page calls the knobs; the field calls them parameters, the word from the section before.
Figure 3 draws all of this at a size that fits on a page: two input words of two numbers each, three units in the middle and five candidate words. The paper’s own example has a vocabulary of 17,964 words, sixty units and a hundred numbers per word, and would be the same drawing at a larger size. The example here comes from the paragraph: “it was” occurs in it ten times, and “the” follows all ten, so nobody chose the right answer. Step through the figure, then let the text set the knobs.
Let’s follow one unit all the way through, so that the other two can be taken on trust. The words “it” and “was” arrive as 0.9, −0.4, −0.3 and 0.8, starting values that nothing has adjusted yet. The topmost unit multiplies them by its own four numbers, 0.3, −0.1, −0.6 and −0.6 (the four lines running into that unit in the drawing), which gives 0.27, 0.04, 0.18 and −0.48. Added together they come to 0.01; the unit adds its own 0.2, making 0.21; and squashing 0.21 leaves it almost unchanged, which is the 0.2 printed beside that unit. The other two units do the same with their own numbers and come out at 0.4 and −0.9.
- step 1 / numberseach of the two words is already a pair of numbers, which is what the section before this one bought. it, was → 0.9, -0.4, -0.3, 0.8.
- step 2 / hiddenevery hidden unit adds up all four, adds its own offset, and squashes the answer between -1 and 1. the three come out 0.2, 0.4, -0.9.
- step 3 / shareseach candidate word scores the three hidden values, and the scores are turned into shares that add to 100. “was” leads on 33.9 per cent.
surprise at “the”: 2.9
The last step turns the scores into shares that add up to one. This page calls that step the sharing box, and every later calculation on the page uses it:
read: every score is made positive, and the scores are scaled so they add to one.
The exponential makes every score positive, and dividing by the total makes the shares add to one. The field’s name for this step is the softmax. The calculation gives every candidate word a share of its belief and leaves the choice of word to a later step. This is the same kind of answer the counting table read off its tally, computed this time instead of looked up.
With the starting knobs the answer is wrong. “Was” leads on 33.9 per cent, while “the”, the word that actually follows, comes last on 5.7. That is expected: if the answer came out right before anything had set the knobs, the knobs would not be doing anything.
Setting the knobs#
Nobody sets the knobs by hand: there are far too many of them, and nobody knows what values they should have. They are worked out from text instead, a small step at a time.
Take a piece of text where you already know which word came next, run it through the calculation and look at the shares that come out. The right word will have received some share, and you want that share to be larger, so you measure how far short it fell. That gap is a single number, which figure 3 calls surprise and the field calls loss. Now work out, for each knob in turn, which way you would have to turn it to make that number a little smaller, and turn every knob a little that way. Then do the same again on the next piece of text, and keep going.
The set of directions you work out is called the gradient, and the arithmetic that carries it back through a stack of layers is called backpropagation. You do not need the details of either to follow the rest of this page, only the fact that the knobs are set by the text rather than by a person.
The nudge button in figure 3 performs one of those steps, on this one example. Two presses take “the” from last place to first, from 5.7 per cent to 37.9, and bring the surprise down from 2.9 to 1.0. The lines in the drawing change width as you press, because the line widths are the knobs. If you keep pressing, the share for “the” keeps growing, but being sure about one sentence is not the same as knowing anything, and the figure says so once the share passes 90 per cent.
However, even with its knobs set, this calculation only ever reads a fixed number of words: figure 3 reads two, and a real one might read ten. Anything before those words is ignored. There is also a subtler problem among the words it does read. Each slot (the position a word occupies in the input) has its own knobs, so how much a word counts depends on the slot it happens to occupy, not on which word it is. The word in the second slot cannot matter more than the word in the fifth just because of what it says.
One answer was to stop using slots at all: feed the words in one at a time and carry a running summary forward, so that any number of words can go in. Everything read so far then has to fit in a summary of a fixed size. The paper that got past this, described in the next section, states the problem in its abstract: the authors conjecture “that the use of a fixed-length vector is a bottleneck”, and their introduction cites a result showing that such a method “deteriorates rapidly as the length of an input sentence increases”. Neither the fixed window nor the fixed-size summary lets the calculation reach back to the one earlier word that settles the next, wherever that word sits.
looking back /
Go back to “it was the age of ___”. The word that settles the ending is “age”. A table with one word of memory cannot see it, and a fixed window helps little, because the window can only treat “age” as whatever happens to be sitting in the fourth slot. What is needed is for the word being worked out to look back at each earlier word and judge how much that word bears on what comes next, based on the words themselves rather than on where they sit.
The mechanism that does this is called attention, and it came from machine translation. In the 2015 paper that introduced it, by Bahdanau, Cho and Bengio, the translation system searches the source sentence every time it produces a word, looking for “a set of positions … where the most relevant information is concentrated”, instead of working from one summary of the whole sentence.
The version used on the rest of this page gives every word three separate vectors, and figure 4 prints all three. The first is what the word is asking for, the second is what the word offers to anything doing the asking, and the third is what the word hands over if it is chosen. The 2017 paper “Attention Is All You Need” calls these the query, the key and the value, and describes the whole step as mapping “a query and a set of key-value pairs to an output”, where the output is a weighted sum of the values and the weight on each value comes from comparing the query with that value’s key.
Figure 4 runs one copy of this step (called a head; a real model runs many side by side) on the six words of “it was the age of wisdom”. Each word carries three vectors of two numbers, chosen by hand so the arithmetic can be read; the scoring, the weighting and the mixing are the real formula. Pick a word to compute its new value, and the three steps fill in as they run.
- it
- was
- the
- age
- of
- wisdom
computing the new value for of · asks for (2, 1)
| word | offers | score | weight | hands over |
|---|---|---|---|---|
| it | (0, 1) | 0.71 | 0.02 | (0, 1) |
| was | (1, 0) | 1.41 | 0.05 | (1, 1) |
| the | (1, 1) | 2.12 | 0.09 | (1, 0) |
| age | (3, 0) | 4.24 | 0.79 | (5, 2) |
| of | (0, 2) | 1.41 | 0.05 | (0, 2) |
| wisdom | · | · | · | · |
- step 1 / scoreeach word offers a vector; the two multiply and add. “of” against “age”: 2×3 + 1×0 = 6, then ÷ √2 = 4.24.
- step 2 / weighthe scores become weights that add to 1.00. “age” takes 0.79 of the mix.
- step 3 / mixeach word hands over its value, weighted. the new value for “of” is (4.09, 1.74).
every word’s weights at once · rows ask, columns answer · the blank half is what a word may not look at
| it | was | the | age | of | wisdom | |
|---|---|---|---|---|---|---|
| it | 1.00 | |||||
| was | 0.80 | 0.20 | ||||
| the | 0.20 | 0.40 | 0.40 | |||
| age | 0.12 | 0.12 | 0.25 | 0.51 | ||
| of | 0.02 | 0.05 | 0.09 | 0.79 | 0.05 | |
| wisdom | 0.01 | 0.01 | 0.04 | 0.74 | 0.02 | 0.18 |
read: each word scores every earlier word, the scores become weights that add to one, and the word’s new value is the weighted mix.
Here is one comparison all the way through. The word being worked out is “of”, and what “of” asks for is the pair of numbers (2, 1). What “age” offers is (3, 0). Multiply the two pairs together position by position and add the results: 2 × 3 is 6 and 1 × 0 is 0, so the score is 6. Then divide by the square root of how many numbers are in a vector. Each vector here has two, so the divisor is √2, about 1.41, and 6 divided by 1.41 gives 4.24, the number in the figure’s score column.
Repeating this against every earlier word, and against “of” itself, gives five scores: 0.71 for “it”, 1.41 for “was”, 2.12 for “the”, 4.24 for “age” and 1.41 for “of”. The sharing box from the previous section turns these five scores into weights that add to one: “age” receives 0.79, and the other four words share the remaining 0.21. Finally, each word hands over its third vector multiplied by its weight, and the five results are added together. “Age” hands over (5, 2) with a weight of 0.79, which is why the new value for “of” comes out at (4.09, 1.74), much closer to what “age” hands over than to anything else in the sentence.
The paper does not claim to have proved that the division is needed, and gives its reason as a suspicion: when vectors are long, the products grow large and push the sharing box into a region where the numbers barely move during training, and dividing keeps the scores in a range where the sharing box still responds.
The blank entries in figure 4 are deliberate. A word is not allowed to look at anything that has not been written yet, so that a prediction “can depend only on the known outputs at positions less than” the one being predicted. Without this restriction, a model asked to continue a sentence could read the answer off the end of the sentence.
The 0.79 is what a table of counts could not produce. Working out “of”, the calculation takes most of its answer from “age”, which sits two places before the word to be predicted, where the one-word table could not see it. The calculation reached “age” because of what the word is, not because of where it sits.
However, attention has no information about the order of the words. The score one word gives another depends only on their vectors, so “of” would give “age” the same score whether “age” sat just before it or five words back. In the 2017 design the position is therefore added into the numbers before attention runs, as an extra ingredient of the input.2
Comparing every word with every earlier word also has a cost, which the 2017 paper gives in a table: the work in one attention layer is proportional to the square of the length of the text, so doubling the length quadruples the work. This is the explosion from when the table explodes in a new form. There, the limit on how far back a model could see was a table nobody could fill; here, it is an amount of arithmetic somebody has to pay for.
the stack /
A language model is mostly built from calculations this page has already described. Take the attention step from looking back, put a layer of knobs like the one in words as numbers after it, and you have a pair of layers. Stack that pair on top of itself a dozen times and you have a language model. The 2018 paper “Improving Language Understanding by Generative Pre-Training” built one this way and describes it in a single sentence: a stack that “applies a multi-headed self-attention operation over the input context tokens followed by position-wise feedforward layers to produce an output distribution over target tokens”. The model in that paper has twelve of these layers.
“Attention Is All You Need”, the 2017 paper that figure 4 follows, built two of these stacks: one to read a sentence in one language and one to write it out in another. The paper names the design the transformer. A language model keeps the writing stack, discards the reading stack and is trained on nothing but the next word, over a very large amount of ordinary text. A transformer with only the writing stack is called a decoder-only transformer.
Every word of the input goes up the whole stack at once, and at the top of the stack is the sharing box again: one score for every word in the vocabulary, turned into shares. A set of shares like this, one per word, is what the quoted sentence calls a distribution.
A full-size model has billions of knobs and is trained on trillions of words of text. A model that is large in both ways is called a large language model (LLM). Those billions of knobs are set by the same nudge that figure 3 performs on one example.
What has been measured about size is narrower than the common claim that bigger is better. Kaplan and colleagues report that the loss “scales as a power-law with model size, dataset size, and the amount of compute used for training”, with some trends “spanning more than seven orders of magnitude”. Loss is the number from setting the knobs: how surprised the model is by the word that actually came next, so a bigger model is predictably less surprised by text. However, whether it understands more is a different question, and that measurement does not answer it.
At the top of the stack no word has been chosen yet. There is a share for every word in the vocabulary, and one of them still has to be picked.
picking a word /
The obvious way to pick is to take the word with the largest share. That is what figure 1 did, and it is how the table wrote “it was the age of times” and then wrote the same words again for ever.
In figure 1 the loop was forced, because after “times” the only word that ever occurred in the paragraph was “it”. A model trained on a large amount of text has a share for every word it knows, and it loops anyway. Holtzman and colleagues measured this, and the reason for it: “the probability of a repeated phrase increases with each repetition, creating a positive feedback loop”. In other words, each repetition makes the phrase likelier to appear once more. The same paper reports that always taking the likeliest word produces text that is “bland, incoherent, or gets stuck in repetitive loops”, and states the general finding plainly: “maximization is an inappropriate decoding objective for open-ended text generation”.
The alternative is to draw a word at random in proportion to its share, so that a word holding a fifth of the belief comes up about one time in five. Shannon’s book method, from counting what follows, is a draw of this kind: opening the book at a random page and reading forward hands you each follower as often as it actually followed.
Between the two methods there is a dial, which the field calls the temperature. Turn it down and the leading share grows until it takes everything; turn it up and the shares level out towards each other. Figure 5 applies the dial to one of figure 1’s own distributions, the eleven words that followed “of” in the paragraph, so none of its numbers were invented for the figure. Turn the dial, then draw.
after of · temperature 1.00
- times16.7
- wisdom8.3
- foolishness8.3
- belief8.3
- incredulity8.3
- light8.3
- darkness8.3
- hope8.3
- despair8.3
- its8.3
- comparison8.3
the paragraph’s own counts, unchanged: this is figure 1’s distribution.
draw a word after “of”, in proportion to its share
the same distribution at the two ends of the dial · with no script there is nothing to turn, so both ends are drawn
after of · temperature 0.00
- times100.0
- wisdom0.0
- foolishness0.0
- belief0.0
- incredulity0.0
- light0.0
- darkness0.0
- hope0.0
- despair0.0
- its0.0
- comparison0.0
one word, always: the word with the largest share is taken every time.
after of · temperature 2.00
- times12.4
- wisdom8.8
- foolishness8.8
- belief8.8
- incredulity8.8
- light8.8
- darkness8.8
- hope8.8
- despair8.8
- its8.8
- comparison8.8
flattened: the lead is narrower than the counts make it. Higher still and every word tends towards the same 9.1 per cent.
read: the same scaling as before, with every score divided by the temperature first: below one it sharpens, above one it flattens.
This is the sharing box from words as numbers, except that the temperature divides every score before the sharing box runs. At a temperature of 1 the division changes nothing, and the shares are the paragraph’s own counts: “times” has 16.7 per cent, which is two occurrences out of twelve. Below 1 the lead widens. Above 1 the eleven shares move towards an eleventh each, and once the dial passes 1 the figure names that limit, 9.1 per cent, so that the top of the dial is not mistaken for the end of the effect.
At 0 the draw is no longer random. The top word is taken every time, which is where this section started: after “it was the age of” the answer is always “times”, as it was in figure 1 back in counting what follows. The usual middle course is to trim the unlikely words before drawing: keep the few likeliest words, or the smallest set of words whose shares add up to some threshold, and draw from those. The second method is called nucleus sampling, and it comes from the same paper that measured the loops.
Wherever the dial is set, the choice of word is made outside the model. The model produces the shares and stops there. The dial and the draw are a separate procedure that somebody chose, and changing that procedure changes the answer without any number inside the model changing.
from predictor to assistant /
Everything on this page so far continues text from where it stopped: given “it was the age of”, the model returns a word to put next. An assistant behaves differently. When you type a question, the assistant answers the question instead of adding a plausible next sentence to it, so something more has been done to the model. A method for doing that has been published.
Ouyang and colleagues, who published the method in 2022, state the problem in the terms this page has been using. Training on “predicting the next token on a webpage from the internet” is a different goal from “follow the user’s instructions helpfully and safely”, and they call the first goal misaligned. Their method has three steps, and each one is the same nudge as in figure 3, aimed at a different target. Nudging an already trained model on new examples is called fine-tuning.
First, people write out good answers to real questions by hand, and the knobs are nudged towards producing those answers. Second, people compare the model’s own answers to the same question and say which they prefer, and a separate set of knobs is trained to predict which answer a person would pick. Third, the original knobs are nudged towards answers that the second set scores highly. In this way the second set of knobs stands in for the people, scoring answers nobody rated by hand. The first step is called instruction tuning, and the second and third together are called reinforcement learning from human feedback (RLHF).
Their evaluation shows what the tuning is worth against size. Answers from the tuned model with 1.3 billion knobs were preferred to answers from the untuned model with 175 billion, although the tuned model had a hundred times fewer knobs.
This was written in August 2026. The assistants most people have used by now are ChatGPT, Claude, Gemini and Copilot, and none of the companies behind them publishes enough about what they ship for this page to say how any one of them was built. What this page can describe is the published method above.
The tuning does not change the kind of answer the model gives. A tuned model is still the stack, with its knobs set differently: it returns a share for every word it knows, and a procedure outside the model still has to draw one of them. The last section is about what that means for the answer.
sampling versus entailment /
When an assistant answers you, the model gives every word it knows a share of its belief as the next word, and what you typed decides those shares. One word is drawn according to them, then the next word, then the one after that. The answer is a sample, and it does not follow from what you typed: nothing in the process derives it from your input, and no amount of further training changes what kind of thing a sample is.
The model can be made to repeat itself exactly. Turn the dial in figure 5 down to zero and the output becomes predictable: the same input gives the same answer every time. However, at zero the model takes the word with the largest share, and taking the likeliest word is still not deriving it from what you typed. A hash function is just as predictable, and its output is not a conclusion drawn from its input either.
Formal concept analysis , which another article on this site explains, is a method of the opposite kind. There, every concept is a closure of the table it came from: nothing appears in the output that the input does not support, and two people starting from the same table finish with the same lattice. That property is called entailment, and it is what the research page claims for formal concept analysis.
Neither kind is better than the other; they answer different questions. A lattice cannot write you a paragraph, and a language model cannot promise that what it wrote is supported by the data it was given. A bigger model does not change that, since a bigger model still samples.
That promise has to come from outside the model, from a checker of the other kind: a symbolic procedure that tests whether a sentence is supported by the data, and gives everyone the same answer. The airlines table in the formal concept analysis article is a small checker of this kind. Asked whether every airline that serves Africa also serves the Middle East, it answers yes, and the answer can be checked row by row. Research on combining learned models like the one on this page with symbolic reasoning of this kind is called neurosymbolic AI.
However, the table can only check sentences made of its own rows and columns. If a model writes that Austrian Airlines flies to Cairo, that sentence has to be rewritten as “Austrian Airlines serves Africa” before the table can answer. The check itself is exact, and the rewriting is where mistakes get in: rules written by hand recognize only the phrasings someone anticipated, and a second language model doing the rewriting would be sampling too.
notes /
- This page says “word” throughout, where the field says “token”. A token is a piece of a word, chosen by a procedure that starts from single characters and repeatedly merges the commonest neighbouring pair until it has made a set number of merges (forty thousand, in the 2018 paper). Splitting words this way lets a model spell out a word it has never seen, instead of having no entry for it. Nothing on this page changes if you read “token” for “word”: the counting, the vectors, the attention and the draw all work on whatever the pieces are. ↩
- The 2017 paper puts order back by adding a position pattern to each word’s numbers before the first attention step, so that the same word in two places arrives as two different sets of numbers. The paper gives its own choice of pattern and notes that there are “many choices of positional encodings, learned and fixed”. In that design, order is an ingredient of the input, and the attention step itself has no notion of it. ↩
sources /
-
The corpus every figure on this page draws on: the opening paragraph, taken verbatim from the plain text edition and committed with its checksum.
Charles Dickens · 1859 · read as Project Gutenberg eBook #98
-
A Mathematical Theory of Communication
Where this page’s counting comes from: its third section prints English built from word transition probabilities, and the sample Shannon produced by hand out of a book.
C. E. Shannon · The Bell System Technical Journal · 1948 · 27, 379–423 and 623–656 · read as the reprint with corrections
-
Prediction and Entropy of Printed English
The experiment this page opens on: people guessing a text one letter at a time, and the sentence about completing an unfinished phrase in conversation.
C. E. Shannon · The Bell System Technical Journal · 1951 · 30(1), 50–64
-
A Neural Probabilistic Language Model
The turn from looking a prediction up to computing it, in one paper (words as learned vectors), and the page’s source for the curse of dimensionality and for the cat and the dog.
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, Christian Jauvin · Journal of Machine Learning Research · 2003 · 3, 1137–1155
-
Neural Machine Translation by Jointly Learning to Align and Translate
Where attention comes from, in translation rather than here, and where the fixed-size summary it was invented to get past is named as the bottleneck.
Dzmitry Bahdanau, KyungHyun Cho, Yoshua Bengio · ICLR 2015 · arXiv:1409.0473
-
The attention box, its three roles and its masking, the table that prices attention at the square of the length, and the name “transformer”; the note on order points at its section on position.
Ashish Vaswani and seven co-authors · NIPS 2017
-
Improving Language Understanding by Generative Pre-Training
The shape of the model: one stack of the 2017 paper’s two, trained on nothing but the next word, in the sentence quoted from it.
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever · technical report, OpenAI · 2018
-
Training Compute-Optimal Large Language Models
The scale of the training text: a model with seventy billion knobs, trained on 1.4 trillion tokens.
Jordan Hoffmann and twenty-one co-authors · NeurIPS 2022
-
Scaling Laws for Neural Language Models
The page’s one claim about size: the loss falls predictably as the model, the data and the training compute grow.
Jared Kaplan and nine co-authors · arXiv:2001.08361 · 2020
-
The Curious Case of Neural Text Degeneration
Why always taking the likeliest word fails, with the feedback loop behind it measured, and where nucleus sampling is proposed.
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, Yejin Choi · ICLR 2020 · arXiv:1904.09751
-
Training language models to follow instructions with human feedback
The three tuning steps, described by one group that ran all three, and the comparison that shows what tuning is worth against a model a hundred times the size.
Long Ouyang and nineteen co-authors · NeurIPS 2022
-
Neurosymbolic AI: the 3rd wave
The field’s name and its scope: learning in neural networks combined with symbolic reasoning, including a network that hands its output to a symbolic system.
Artur d’Avila Garcez, Luís C. Lamb · Artificial Intelligence Review · 2023 · 56, 12387–12406
-
Speech and Language Processing
The standard reference: the definition of a large language model, and a pointer to n-grams, smoothing and most of what this page mentions and does not teach.
Daniel Jurafsky, James H. Martin · third edition draft of 19 August 2026
more /
Something wrong on this page, or something worth adding? Reach me by email.