Prequel · LSTMs

How an LSTM Remembers

“Before attention, models carried memory along a belt, word by word. This is the one to walk before stop 3 if you want the full story.”

Unpacking the world: 1.7 MB of drawings and measured data.

AgenticAmit A field guide to LSTM networks

Field guide · LSTM networks

How a network remembers ten words back.

A model reading one word at a time has to carry something from the start of a sentence to the end, past words that point the wrong way. An LSTM does it with a belt, three gates, and one plus sign.

Scroll to walk the sentence ↓

Fig. 01 · the test sentenceTACL 2016
Linzen et al. example (2): The keys to the cabinet are on the table. Intervening nouns with the opposite number from the subject are called agreement attractors.

Linzen, Dupoux & Goldberg · Assessing the ability of LSTMs…

How does it know it's "are"?

Pinned · the forget gateNeural Computation 2000
Gers, Schmidhuber and Cummins: Our remedy is a novel, adaptive forget gate that enables an LSTM cell to learn to reset itself at appropriate times.

Gers, Schmidhuber & Cummins · Learning to Forget

Pinned · the whole cellPyTorch docs
PyTorch nn.LSTM equations: i, f, g, o gates; c_t = f_t ⊙ c_{t-1} + i_t ⊙ g_t; h_t = o_t ⊙ tanh(c_t).

torch.nn.LSTM

Pinned · the original ideaNeural Computation 1997
Hochreiter and Schmidhuber: LSTM can learn to bridge minimal time lags in excess of 1000 discrete time steps by enforcing constant error flow through constant error carrousels within special units.

Hochreiter & Schmidhuber · Long Short-Term Memory

Pinned · the problem, namedIEEE TNN 1994
Bengio, Simard and Frasconi: gradient based learning algorithms face an increasingly difficult problem as the duration of the dependencies to be captured increases.

Bengio, Simard & Frasconi · Learning long-term dependencies…

Pinned · where it wentNeurIPS 2024
xLSTM: How far do we get in language modeling when scaling LSTMs to billions of parameters?

Beck et al. · xLSTM: Extended Long Short-Term Memory

Lane 2LSTM

A belt and three gates

The belt is the cell state. We follow cell 4 of 8.

Lane 1plain RNN

One note, rewritten

Every word, the whole note is rewritten.

Output · LSTMreads the belt
are

p(are) = 1.000 · right

Output · RNNreads its note
is

p(are) = 0.464 · wrong

Fig. 05 · our sweep3 seeds per box
Which distances each network learned

Filled dot: that training run got at least 95% of held-out sentences right, with every distractor noun disagreeing. Hidden size 8, 3,000 steps.

Fig. 06 · cell 4, two subjectssame network
It keeps a tally, not a word

Olive: "the keys to the cabinet…" ends at +3.56. Clay: "the cabinet to the table…" ends at −5.96.

Start with a sentence

You already know the answer.

"The keys to the cabinet by the door near the table ___." Is or are? It's "are". To get there you carried "keys" across nine words and ignored three singular nouns that point the wrong way.

A network reading one word at a time has to do the same. Researchers use exactly this kind of sentence to test them.

Lane 1 · a plain RNN

Read a word, rewrite the note.

A recurrent neural network (RNN) reads one word at a time and carries a small note forward: a list of numbers called the hidden state. At every word it rewrites the whole note from the old note plus the new word.

Lane 1 · the run

It ends on a coin flip.

We trained this RNN on thousands of sentences like ours. Watch its guess for "are": it hovers around 50% at every word, and on the last one it says "is". On held-out sentences it got 37% right.

The catch

The lesson never reaches "keys".

A network learns by sending an error signal backward from the answer to every word that caused it. In an RNN that signal gets multiplied down at every step. From three phrases away, 0.03% of it reaches "keys". It can't learn from what it can't feel.

MeasuredFresh networks, 5 seeds × 100 sentences each.

Lane 2 · the fix

Add a belt.

An LSTM (long short-term memory) keeps the RNN's note and adds a second one that runs straight through every word, like a conveyor belt: the cell state. Nothing rewrites the belt wholesale. At each word, three gates decide what happens to it.

Gate 1 · forget

How much of the belt to keep.

The forget gate is a number from 0 to 1 that multiplies what's already on the belt. 1 keeps everything; 0 wipes it. At "keys" this cell's forget gate dips to 0.81. Everywhere from "the" to "the" after it, it sits at 0.97 or higher.

2000This gate came later: Gers, Schmidhuber and Cummins added "a novel, adaptive 'forget gate'" to the original LSTM.

Gate 2 · input

What to write, and how much.

The word proposes something to add, called the candidate (between −1 and 1). The input gate decides how much of it gets through. At "keys" the gate is nearly shut, 0.17, so this cell barely changes.

Gate 3 · output

What to read off, right now.

The output gate decides how much of the belt shows up in the note the rest of the network sees at this word. The belt can hold something without saying it yet.

The one line that matters

Keep, then add.

c = f × c + i × g

Multiply the belt by the forget gate, then add the gated candidate. Adding instead of rewriting is the whole trick: what's on the belt survives unless a gate chooses to let it go.

Words 3 and 4

The tally starts to lean.

Here's something the textbook picture skips. In our trained network "keys" doesn't stamp "plural" onto the belt in one go. The next words, read in its light, push this cell up: +0.41, then +0.96.

Word 5 · the distractor

"Cabinet" barely moves it.

A singular noun arrives. The input gate opens to 0.65, but the candidate is only +0.15, so it adds about +0.10. The forget gate holds at 0.99. The network learned that a noun after "to the" isn't the one that decides the verb.

Words 6 to 11

Door, table: same story.

Two more singular nouns, and the forget gate never drops below 0.99 until the last word. The tally keeps climbing, to +3.56.

The verb

One says "are". One says "is".

The LSTM reads the belt through its output gate and says "are", with 100% confidence. The RNN, trained on exactly the same sentences, says "is".

Why it could learn this

The belt carries the lesson back too.

The belt that carries memory forward also carries the error signal backward, and it's only multiplied by forget gates. From three phrases away, 7.3% of the signal reaches "keys", against 0.03% for the RNN. That's about 240 times more.

1997Hochreiter and Schmidhuber called this "constant error flow through constant error carrousels".

But here's the catch

Longer, not forever.

We trained both networks from scratch at five distances, three times each. The LSTM learned every run up to 4 phrases, and 2 of 3 at 5–6. At 7 or more, neither learned at this size. The belt stretches memory; it doesn't make it infinite.

What the cell really stores

A tally, not a word.

Give the same cell a singular subject ("the cabinet to the table…") and it runs the other way, down to −5.96. It isn't storing the word "keys". It keeps a running lean, and every word nudges it.

Where LSTMs went

Transformers took over. The idea didn't leave.

Transformers now do most language work by looking at every word at once instead of carrying a belt. In 2024 a team including Sepp Hochreiter, the LSTM's co-inventor, asked how far LSTMs go when scaled to billions of parameters, and introduced xLSTM.

Take this with you

An LSTM doesn't remember more. It forgets on purpose.

Three gates and a plus sign: keep, write, read. The next time a model gets a long sentence right, you'll know what had to survive the trip, and what it chose to let go.

Amit, thumbs up