Business Pill 23 · How a model learns before it has a job
Before it has any job, a model learns by playing one simple game trillions of times: guess the next word. That stage is called pretraining.
A short explainer video. The numbers in it are round numbers for illustration.
The Short Answer
Before anyone teaches it a task, a model plays a simple game. Take a sentence, hide the next word, and ask the model to guess it. Then show the real word and adjust the model a little, so its next guess is better.
Trillions of Guesses
Now repeat that game trillions of times, across books, websites, code and conversations. Nobody labels anything. The text itself provides the answer. That is what makes it possible at this scale.
To guess well, the model has to pick up a great deal along the way: grammar, facts, how an argument is built, how code works. None of it is taught directly. It all comes from predicting the next word.
The Expensive Stage
This stage is called pretraining. It produces a general model, one that knows a lot but has no particular job yet. It is also the expensive stage: thousands of chips running for months. Very few organizations can afford it.
But it is done once, and the result can be reused for thousands of different purposes. This shapes the whole industry: a handful of companies build the general models, and everyone else builds on top of them.
Why It Matters
A model only knows what was in its training text, up to the day that text was collected.
Three Questions to Ask
- What text was the model trained on?
- When does its knowledge end?
- What does it still need to learn for our work?
More Business Pills
- The Memory Wall: Why a Faster AI Chip Is Not Faster AI
- Tokens per Watt: What an AI Data Centre Actually Produces
- Why AI Can’t Be Both Instant and Cheap: Latency vs Throughput
- The Model and the Harness: Why Same-Model Products Differ
- Context, Not Capability: Why a Smart AI Model Gives Poor Answers
- Discardable Software: When Code Is Cheap Enough to Throw Away
- Human in the Loop vs Human on the Loop: Supervising AI Agents
- The AI Audit Problem: When Making Work Is Cheaper Than Checking It
- AI Evaluation as Acceptance Test: How to Know It’s Good Enough
- Extensibility Is Control: Who Holds the Power in an AI Product
- RLHF Explained: How an AI Model Learns What People Prefer
- Fine-Tuning Explained: How to Adapt a General AI Model to One Job
- Tokens Explained: The Unit AI Reads, Writes and Bills In
- Embeddings Explained: How a Machine Compares Meaning
- RAG Explained: How a Model Answers From Your Documents
- AI Hallucination Explained: Why Models State False Things
- Distillation Explained: How a Small Model Learns From a Large One
- Reasoning Models Explained: What Changes When AI Thinks First
- Tool Use Explained: How an AI Model Goes From Text to Action
- Capex and Depreciation Explained: Why Chip Lifetime Drives AI Profits
- Run-Rate Revenue Explained: What an AI Company’s Number Means
- Backlog Explained: Revenue That Is Signed but Not Yet Earned
- Switching Costs Explained: Why It Is Hard to Leave an AI Supplier









