Season one · stop 1 · Tokenization

How a Sentence Gets Chopped

“Before a model reads a single word, your sentence gets cut into pieces. We'll watch it happen, and you'll see why counting the r's in strawberry is harder than it looks.”

Unpacking the world: 2.5 MB of drawings and measured data.

AgenticAmit A field guide to tokenization

Field guide · tokenization

How a sentence gets chopped.

Before a chatbot reads a single word, a press cuts your text into pieces and swaps each piece for a number. Here's the press, the books that trained it, and what it does to one question.

Scroll to run the press ↓

The strip · your question, one block per character31 characters
Pinned · start from bytesOpenAI 2019
Radford et al.: In contrast, a byte-level version of BPE only requires a base vocabulary of size 256. Radford et al.: Since our approach can assign a probability to any Unicode string, this allows us to evaluate our LMs on any dataset regardless of pre-processing, tokenization, or vocab size.

Radford et al. · Language Models are Unsupervised Multitask Learners

The book · what our press readProject Gutenberg

+ the rest of Alice's Adventures in Wonderland, Pride and Prejudice, Frankenstein and The Adventures of Sherlock Holmes

1,886,394 bytes of text, every neighbouring pair counted
The rule book · merges, in the order learnedpair · times seen
Glue the most common pair. Repeat.

    … 7,984 more rules. Rule 6,400 glues “␣st” and “raw” into “␣straw”: seen 9 times in four books.

    Pinned · the whole algorithmACL 2016
    Sennrich et al., section 3.2: Byte Pair Encoding (BPE) (Gage, 1994) is a simple data compression technique that iteratively replaces the most frequent pair of bytes in a sequence with a single, unused byte. We adapt this algorithm for word segmentation. Instead of merging frequent pairs of bytes, we merge characters or character sequences. Sennrich et al., Algorithm 1: Learn BPE operations, a short Python listing.

    Sennrich, Haddow & Birch · Neural Machine Translation of Rare Words with Subword Units

    that's the whole press!

    Pinned · where the trick came fromFebruary 1994
    A New Algorithm for Data Compression, Philip Gage. Gage: This article describes a simple general-purpose data compression algorithm, called Byte Pair Encoding (BPE), which provides almost as much compression as the popular Lempel, Ziv, and Welch (LZW) method.

    Philip Gage · A New Algorithm for Data Compression · Dr. Dobb's archive

    Pinned · the real pressGitHub
    tiktoken is a fast BPE tokeniser for use with OpenAI's models.

    openai/tiktoken

    What the model receivesGPT-4o's press

    Eight numbers. Count the r's.

    Fig. 04 · one word, three waysGPT-4o's press
    Fig. 05 · how many pieces each press knowsvocabulary
    Fig. 06 · the price boardUDHR Article 1 · pieces per press
    One sentence, many languages
    GPT-4o's pressGPT-2's pressraw UTF-8 bytes
    Pinned · the unfairness, measuredNeurIPS 2023
    Petrov et al.: we show how disparity in the treatment of different languages arises at the tokenization stage, well before a model is even invoked. The same text translated into different languages can have drastically different tokenization lengths, with differences up to 15 times in some cases.

    Petrov, La Malfa, Torr & Bibi · Language Model Tokenizers Introduce Unfairness Between Languages

    Pinned · a model with no pressACL 2025
    Pagnoni et al.: We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency and robustness. BLT encodes bytes into dynamically sized patches, which serve as the primary units of computation.

    Pagnoni et al. · Byte Latent Transformer: Patches Scale Better Than Tokens

    Start with a question

    Count the r's.

    "How many r's are in strawberry?" You can see the answer: three, right there in the word. That's because you're reading letters.

    Let's follow this exact sentence into a model and see whether the letters survive the trip.

    Step 1 · bytes

    Every letter is already a number.

    Computers store text as bytes, numbers from 0 to 255. In the standard encoding (UTF-8), "r" is byte 114. Our 31-character question is 31 bytes.

    That's where the press starts: 256 possible pieces, one per byte value. Any text in any language can be written with them.

    Step 2 · the book

    The press learns by reading.

    Before it can cut anything, the press reads a pile of text and counts every pair of neighbouring pieces. We built our own and fed it four old novels from Project Gutenberg: 1.9 million bytes.

    The most common pair was a space followed by "t". In this one sentence of Alice it shows up 8 times. Across all four books: 40,947.

    Step 3 · the rules

    Glue the most common pair. Repeat.

    The press glues that pair into one new piece, gives it the next free number (256), and counts again. Then "h" + "e". Then a space + "a". Each glue becomes a rule, kept in order.

    That's the whole algorithm. It's called byte-pair encoding (BPE). We let ours learn 8,000 rules.

    1994 → 2016It started as a file-compression trick by Philip Gage. Sennrich, Haddow and Birch adapted it to split rare words for translation, and printed the whole thing as a short Python listing.

    The trays · 10 rules

    The first cuts are tiny.

    Now run our question through the rules, in order. After the first 10, 31 pieces become 27: a space glues onto "a", and "r" + "e" makes "re". (␣ marks a space.)

    Look under each piece: the number is what gets passed on. New pieces get numbers from 256 up.

    The trays · 100 rules

    Common chunks go first.

    After 100 rules: 22 pieces. "␣in" is one piece now, and so is "er", which was rule 14. The chunks English uses most get glued long before rare ones.

    The trays · 500 rules

    Every r is inside a bigger piece.

    After 500 rules: 16 pieces. The last lone r in strawberry just got glued to its y ("ry", rule 275). The letters are still printed on the blocks, but the model won't see the blocks. It sees the numbers under them.

    The trays · 8,000 rules

    It only learns what it read.

    After all 8,000 rules: 11 pieces. Our press still cuts "␣straw" | "ber" | "ry". Why? The word "strawberry" doesn't appear once in those four books. The press builds its pieces out of whatever it read most.

    The real press

    In GPT-4o's press, it's one piece.

    Now the same question through three real tokenizers, measured with OpenAI's open-source tiktoken. They learned from far more text than four novels. All three cut it into 8 pieces, and in every one, "␣strawberry" (with the space in front) is a single piece.

    In GPT-4o's press, the whole word is number 101830.

    The catch

    Now count the r's.

    This is the whole question as the model gets it: eight numbers. There's no "r" in 101830. Nothing in it says s-t-r-a-w-b-e-r-r-y.

    To count letters, a model has to have picked up, somewhere in training, how each of its pieces is spelled. It can learn that. It just isn't in what it's handed.

    One space changes it

    Same word, different pieces.

    The press cuts the exact bytes, spaces and capitals included. "␣strawberry", with a space in front, is one piece. "strawberry" with no space is three: "st", "raw", "berry". "Strawberry" with a capital is three again, starting with a different number.

    The drawer

    Bigger reading, bigger drawer.

    Every press has a fixed set of pieces, called its vocabulary. Ours knows 8,256. GPT-2's knows 50,257, GPT-4's 100,277 and GPT-4o's 200,019. More pieces means more whole words, so the same text takes fewer.

    But here's the catch

    The same sentence costs more in some languages.

    Here's one sentence, Article 1 of the Universal Declaration of Human Rights, in twelve official translations, run through GPT-4o's press. English takes 33 pieces. Hindi takes 54, Tamil 76, Burmese 134, and Amharic 206: over six times as many.

    Translations aren't word for word, so small gaps don't mean much. Six times does. You pay per piece, and a model's memory is counted in pieces too.

    Why it happens

    The press read some languages far more.

    Rules come from the training text. A language the press saw less gets fewer rules, so it's cut into smaller pieces. Some scripts also start bigger: a Hindi letter takes 3 bytes in UTF-8, where an English letter takes 1.

    The grey bars are raw bytes, the most pieces any press could produce. Hindi's sentence is 499 bytes; English's is 170.

    Newer presses

    The gap shrank. It didn't close.

    The clay bars are GPT-2's press. It cut the Tamil sentence into 664 pieces for 664 bytes: it had never learned a single Tamil rule, so every byte stood alone. GPT-4o's press gets it down to 76. Amharic only went from 309 to 206, while English sat at 33 in all three.

    2023Petrov and colleagues found this happens “at the tokenization stage, well before a model is even invoked”, and the same text showing “differences up to 15 times in some cases”.

    Where it's going

    What if there's no press?

    Researchers at Meta and the University of Washington built a model that reads raw bytes and groups them on the fly, with no fixed vocabulary: the Byte Latent Transformer (ACL 2025). It's a research model. Almost every chatbot you use today still has a press in front of it.

    Take this with you

    A model never reads your words. It reads the pieces.

    A press, trained on somebody else's reading, decides where your sentence gets cut. Next time a chatbot trips over spelling, or the same message costs more in another language, you'll know where it happened: before the model saw anything at all.

    Amit, thumbs up