Season one · stop 2 · Next-token prediction

Learning to Write

“The whole trick behind a chatbot is guessing the next piece, over and over. We trained a tiny model from scratch, so you can walk down its real learning curve and read what it wrote at every stage.”

Unpacking the world: 2.2 MB of drawings and measured data.

AgenticAmit A field guide to how language models learn

Field guide · next-token prediction

Learning to write, one guess at a time.

Every chatbot starts as a machine that plays one game: guess what comes next. We trained a small one from scratch on four old novels and kept what it wrote at every stage. This is the whole trip, from noise to sentences.

Scroll to walk down the hill ↓

The gamefrom Alice, chapter 1
Alice was beginning to get very tire?
What comes next? One character.
Pinned · the same game, with peopleBell System Technical Journal 1951
Shannon 1951: Select a short passage unfamiliar to the person who is to do the predicting. He is then asked to guess the first letter in the passage. If the guess is correct he is so informed, and proceeds to guess the second letter. If not, he is told the correct first letter and proceeds to his next guess. Shannon 1951, example: THERE IS NO REVERSE ON A MOTORCYCLE A FRIEND OF MINE FOUND THIS OUT RATHER DRAMATICALLY THE OTHER DAY, with the number of guesses needed for each letter underneath.

Claude E. Shannon · Prediction and Entropy of Printed English

The guess · top 8 of 99 charactersstep 0
Before any training

The scorecross-entropy loss

Writing · the finished model, one letter at a timebar = how sure it was
Every letter is a guess it then keeps.

Confident, and made up

Pinned · where it startsarXiv 2025
Kalai et al., section 1.1 Errors caused by pretraining: During pretraining, a base model learns the distribution of language in a large text corpus. We show that, even with error-free training data, the statistical objective minimized during pretraining would lead to a language model that generates errors.

Kalai, Nachum, Vempala & Zhang · Why Language Models Hallucinate

Fig. 04 · is it copying?checked against all four books

Fig. 05 · same game, biggerparameters · training text
The recipe doesn't change. The size does.

Circle area to scale. GPT-2 figures as printed in its paper.

Pinned · the same objectiveOpenAI 2019
Radford et al.: … symbols as the product of conditional probabilities (Jelinek & Mercer, 1980) (Bengio et al., 2003): p(x) = product from i = 1 to n of p(s_n | s_1, …, s_n−1) (1), as printed. Radford et al.: WebText … contains slightly over 8 million documents for a total of 40 GB of text. Radford et al.: Our largest model, GPT-2, is a 1.5B parameter Transformer.

Radford et al. · Language Models are Unsupervised Multitask Learners

Take this with you

It learned what usually comes next. Not what's true.

That one game, played millions of times, is enough to pull spelling, grammar and style out of plain text. It's also why a model can sound completely sure while making things up. Everything a chatbot does later is built on top of this.

Amit, calm and reassuring