Business Pill 25 · The unit AI is measured in
A model does not read words. It reads tokens, and tokens are how AI is priced, measured and limited.
A short explainer video. The numbers in it are round numbers for illustration.
The Short Answer
A model does not read words. It reads tokens. A token is a small piece of text, sometimes a whole word, sometimes part of one. As a rule of thumb, 100 tokens is about 75 words.
Counted in Both Directions
Everything you send to the model is input: your question, your documents, the conversation so far. Everything it writes back is output. And output usually costs several times more than input.
A Worked Example
The video uses round numbers. You send a 10-page document and ask for a one-page summary. That is about 5,000 tokens in and 500 out. If input costs $3 per million tokens and output $15, the request costs about two cents.
Two cents sounds like nothing. But an agent working on a task may make hundreds of such calls, and it rereads its whole history each time. So the tokens, and the bill, grow much faster than the work seems to.

Why It Matters
This is why the token is the basic unit of the AI economy. Models are priced in tokens. Chips are measured in tokens per second. Data centres in tokens per watt.
Tokens also set a limit. A model can only hold a fixed number at once. That is its context window. Long documents and long conversations fill it up.
Three Questions to Ask
- How many tokens does one task really use?
- How many of them are output?
- How does that grow with each extra step?
More Business Pills
- The Memory Wall: Why a Faster AI Chip Is Not Faster AI
- Tokens per Watt: What an AI Data Centre Actually Produces
- Why AI Can’t Be Both Instant and Cheap: Latency vs Throughput
- The Model and the Harness: Why Same-Model Products Differ
- Context, Not Capability: Why a Smart AI Model Gives Poor Answers
- Discardable Software: When Code Is Cheap Enough to Throw Away
- Human in the Loop vs Human on the Loop: Supervising AI Agents
- The AI Audit Problem: When Making Work Is Cheaper Than Checking It
- AI Evaluation as Acceptance Test: How to Know It’s Good Enough
- Extensibility Is Control: Who Holds the Power in an AI Product
- RLHF Explained: How an AI Model Learns What People Prefer
- Pretraining Explained: How an AI Model Learns Before Anyone Teaches It
- Fine-Tuning Explained: How to Adapt a General AI Model to One Job
- Embeddings Explained: How a Machine Compares Meaning
- RAG Explained: How a Model Answers From Your Documents
- AI Hallucination Explained: Why Models State False Things
- Distillation Explained: How a Small Model Learns From a Large One
- Reasoning Models Explained: What Changes When AI Thinks First
- Tool Use Explained: How an AI Model Goes From Text to Action
- Capex and Depreciation Explained: Why Chip Lifetime Drives AI Profits
- Run-Rate Revenue Explained: What an AI Company’s Number Means
- Backlog Explained: Revenue That Is Signed but Not Yet Earned
- Switching Costs Explained: Why It Is Hard to Leave an AI Supplier
- Temperature Explained: The Dial That Sets How Predictable an AI Model Is
- Guardrails Explained: Rules Enforced Around an AI Model, Not Inside It
- Benchmarks Explained: Why a Test Score Is Not Your Own Result
- Agent Memory Explained: How an AI Assistant Remembers You Between Chats
- MCP Explained: The Shared Plug That Connects AI Models to Tools
- Quantization Explained: How a Large AI Model Is Shrunk to Fit
- Parameters Explained: Where an AI Model’s Knowledge Lives
- Scaling Laws Explained: Why AI Labs Keep Building Bigger
- Synthetic Data Explained: When a Model Writes Its Own Training Examples
- Transformers and Attention Explained: How AI Reads Every Word at Once









