Tokens Explained: The Unit AI Reads, Writes and Bills In

Business Pill 25 · The unit AI is measured in

A model does not read words. It reads tokens, and tokens are how AI is priced, measured and limited.

A short explainer video. The numbers in it are round numbers for illustration.

The Short Answer

A model does not read words. It reads tokens. A token is a small piece of text, sometimes a whole word, sometimes part of one. As a rule of thumb, 100 tokens is about 75 words.

Counted in Both Directions

Everything you send to the model is input: your question, your documents, the conversation so far. Everything it writes back is output. And output usually costs several times more than input.

A Worked Example

The video uses round numbers. You send a 10-page document and ask for a one-page summary. That is about 5,000 tokens in and 500 out. If input costs $3 per million tokens and output $15, the request costs about two cents.

Two cents sounds like nothing. But an agent working on a task may make hundreds of such calls, and it rereads its whole history each time. So the tokens, and the bill, grow much faster than the work seems to.

Bar chart: about 5,000 tokens in and 500 tokens out for a 10-page document summarised in one page
Input is ten times the output here, yet output costs several times more per token. Round numbers for illustration.

Why It Matters

This is why the token is the basic unit of the AI economy. Models are priced in tokens. Chips are measured in tokens per second. Data centres in tokens per watt.

Tokens also set a limit. A model can only hold a fixed number at once. That is its context window. Long documents and long conversations fill it up.

Three Questions to Ask

  1. How many tokens does one task really use?
  2. How many of them are output?
  3. How does that grow with each extra step?

More Business Pills

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA