The Memory Wall: Why a Faster AI Chip Is Not Faster AI

The Memory Wall: Why a Faster AI Chip Is Not Faster AI

Business Pill 12 · Why a faster chip is not enough

A faster chip is not always a faster AI. The reason is the memory wall, and it comes down to how fast the model can be delivered to the chip.

A short explainer video. The numbers in it are round numbers for illustration.

The Short Answer

An AI chip has two jobs. It calculates, and it fetches. The numbers that make up the model sit in memory, and they have to be carried to the processor before any calculation can happen.

A model is large. Each time it produces one word, it has to read almost all of itself, which means billions of numbers moved from memory to the processor. So the speed of an answer depends on two things: how fast the processor calculates, and how fast memory can deliver the model to it.

What Memory Bandwidth Is

The delivery speed is called memory bandwidth. Think of it as the width of the road between memory and the processor. A wide road moves a lot of the model per second. A narrow road moves little, however powerful the processor at the far end.

For years, processors got faster much more quickly than memory did. The road became the narrow part. The processor is ready, and it is waiting for data.

A Worked Example

The video uses round numbers. A model takes up 100 GB, and memory can deliver 1,000 GB per second. The model can be read 10 times a second, and since each word needs one read, that is 10 words per second at most. This holds however fast the processor is.

Now double the speed of the processor. Nothing changes: still 10 words per second. Double the memory bandwidth instead and you get 20. The point is which change helps, not the specific numbers.

Bar chart: words per second at most is 10 today, 10 with a processor twice as fast, and 20 with twice the memory bandwidth
Doubling the processor changes nothing in this example. Doubling memory bandwidth doubles the ceiling. Round numbers for illustration.

Why It Matters

Engineers call this the memory wall: the point where the limit is no longer how fast you can compute, but how fast you can feed the computer. The video gives three reasons it matters.

First, fast memory becomes as scarce and as valuable as the processors themselves. Second, smaller models run faster simply because there is less to read. Third, serving many users together helps, because one read of the model can answer many requests at once.

Three Questions to Ask

  1. How large is the model?
  2. How fast can memory feed it?
  3. How many requests share each read of the model?

See It in the News

Fractile’s Case That FLOPs Are the Wrong Metric. A news piece on the same bottleneck: how compute and memory bandwidth have scaled at different rates.

More Business Pills

Business Engineer

The Map of AI

The Business Engineer framework maps where each part of the AI stack sits and who holds the leverage. These pills are short lessons from the same library.

Explore the Map of AI →
Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA