Count the r's.
"How many r's are in strawberry?" You can see the answer: three, right there in the word. That's because you're reading letters.
Let's follow this exact sentence into a model and see whether the letters survive the trip.

Season one · stop 1 · Tokenization
How a Sentence Gets Chopped
“Before a model reads a single word, your sentence gets cut into pieces. We'll watch it happen, and you'll see why counting the r's in strawberry is harder than it looks.”
Unpacking the world: 2.5 MB of drawings and measured data.
Before a chatbot reads a single word, a press cuts your text into pieces and swaps each piece for a number. Here's the press, the books that trained it, and what it does to one question.
Scroll to run the press ↓
Radford et al. · Language Models are Unsupervised Multitask Learners
+ the rest of Alice's Adventures in Wonderland, Pride and Prejudice, Frankenstein and The Adventures of Sherlock Holmes
… 7,984 more rules. Rule 6,400 glues “␣st” and “raw” into “␣straw”: seen 9 times in four books.
Sennrich, Haddow & Birch · Neural Machine Translation of Rare Words with Subword Units
that's the whole press!
Philip Gage · A New Algorithm for Data Compression · Dr. Dobb's archive
openai/tiktoken
Eight numbers. Count the r's.
Petrov, La Malfa, Torr & Bibi · Language Model Tokenizers Introduce Unfairness Between Languages
Pagnoni et al. · Byte Latent Transformer: Patches Scale Better Than Tokens
"How many r's are in strawberry?" You can see the answer: three, right there in the word. That's because you're reading letters.
Let's follow this exact sentence into a model and see whether the letters survive the trip.

Computers store text as bytes, numbers from 0 to 255. In the standard encoding (UTF-8), "r" is byte 114. Our 31-character question is 31 bytes.
That's where the press starts: 256 possible pieces, one per byte value. Any text in any language can be written with them.
Before it can cut anything, the press reads a pile of text and counts every pair of neighbouring pieces. We built our own and fed it four old novels from Project Gutenberg: 1.9 million bytes.
The most common pair was a space followed by "t". In this one sentence of Alice it shows up 8 times. Across all four books: 40,947.
The press glues that pair into one new piece, gives it the next free number (256), and counts again. Then "h" + "e". Then a space + "a". Each glue becomes a rule, kept in order.
That's the whole algorithm. It's called byte-pair encoding (BPE). We let ours learn 8,000 rules.
1994 → 2016It started as a file-compression trick by Philip Gage. Sennrich, Haddow and Birch adapted it to split rare words for translation, and printed the whole thing as a short Python listing.
Now run our question through the rules, in order. After the first 10, 31 pieces become 27: a space glues onto "a", and "r" + "e" makes "re". (␣ marks a space.)
Look under each piece: the number is what gets passed on. New pieces get numbers from 256 up.
After 100 rules: 22 pieces. "␣in" is one piece now, and so is "er", which was rule 14. The chunks English uses most get glued long before rare ones.
After 500 rules: 16 pieces. The last lone r in strawberry just got glued to its y ("ry", rule 275). The letters are still printed on the blocks, but the model won't see the blocks. It sees the numbers under them.
After all 8,000 rules: 11 pieces. Our press still cuts "␣straw" | "ber" | "ry". Why? The word "strawberry" doesn't appear once in those four books. The press builds its pieces out of whatever it read most.
Now the same question through three real tokenizers, measured with OpenAI's open-source tiktoken. They learned from far more text than four novels. All three cut it into 8 pieces, and in every one, "␣strawberry" (with the space in front) is a single piece.
In GPT-4o's press, the whole word is number 101830.

This is the whole question as the model gets it: eight numbers. There's no "r" in 101830. Nothing in it says s-t-r-a-w-b-e-r-r-y.
To count letters, a model has to have picked up, somewhere in training, how each of its pieces is spelled. It can learn that. It just isn't in what it's handed.

The press cuts the exact bytes, spaces and capitals included. "␣strawberry", with a space in front, is one piece. "strawberry" with no space is three: "st", "raw", "berry". "Strawberry" with a capital is three again, starting with a different number.
Every press has a fixed set of pieces, called its vocabulary. Ours knows 8,256. GPT-2's knows 50,257, GPT-4's 100,277 and GPT-4o's 200,019. More pieces means more whole words, so the same text takes fewer.
Here's one sentence, Article 1 of the Universal Declaration of Human Rights, in twelve official translations, run through GPT-4o's press. English takes 33 pieces. Hindi takes 54, Tamil 76, Burmese 134, and Amharic 206: over six times as many.
Translations aren't word for word, so small gaps don't mean much. Six times does. You pay per piece, and a model's memory is counted in pieces too.
Rules come from the training text. A language the press saw less gets fewer rules, so it's cut into smaller pieces. Some scripts also start bigger: a Hindi letter takes 3 bytes in UTF-8, where an English letter takes 1.
The grey bars are raw bytes, the most pieces any press could produce. Hindi's sentence is 499 bytes; English's is 170.

The clay bars are GPT-2's press. It cut the Tamil sentence into 664 pieces for 664 bytes: it had never learned a single Tamil rule, so every byte stood alone. GPT-4o's press gets it down to 76. Amharic only went from 309 to 206, while English sat at 33 in all three.
2023Petrov and colleagues found this happens “at the tokenization stage, well before a model is even invoked”, and the same text showing “differences up to 15 times in some cases”.
Researchers at Meta and the University of Washington built a model that reads raw bytes and groups them on the fly, with no fixed vocabulary: the Byte Latent Transformer (ACL 2025). It's a research model. Almost every chatbot you use today still has a press in front of it.
A press, trained on somebody else's reading, decides where your sentence gets cut. Next time a chatbot trips over spelling, or the same message costs more in another language, you'll know where it happened: before the model saw anything at all.