Business Pill 27 · Answering from your own documents
RAG lets a model answer from your own documents without retraining it: look up the right pages first, then hand them to the model with the question.
A short explainer video. The numbers in it are round numbers for illustration.
The Short Answer
A model knows only what it was trained on. It has never seen your contracts, your manuals, or last week’s prices. Ask about them and it either refuses or guesses.
The fix is simple: before the model answers, look up the relevant pages, then hand them to the model together with the question.
Three Steps
First, retrieve: search your documents for the passages closest to the question. Second, augment: add those passages to the request. Third, generate: the model writes its answer using what it was just given.
An Example
An employee asks, how many days of leave do I get for a new child? The system finds the right page of the staff handbook, the model reads it, and answers with the exact number.
That is retrieval-augmented generation, usually shortened to RAG.
Why It Is Popular, and Its Weak Point
It is popular for good reasons. The knowledge stays current, because you update the documents, not the model. The answer can show its sources. And private information stays in your own systems.
But it has one weak point: the answer is only as good as what was retrieved. If the search brings back the wrong page, the model answers confidently from the wrong page.
Three Questions to Ask
- Are the right documents in the system?
- Does the search find the right passage?
- Can we see which source each answer used?
See It in the News
Cloudflare Web Search API Prices $0.25 to $7.00. A news piece on a service that hands live web results to a model before it answers.
TextQL Ontology GA: Context, Not Capability. A news piece on giving a model the context of a business’s own data.
More Business Pills
- The Memory Wall: Why a Faster AI Chip Is Not Faster AI
- Tokens per Watt: What an AI Data Centre Actually Produces
- Why AI Can’t Be Both Instant and Cheap: Latency vs Throughput
- The Model and the Harness: Why Same-Model Products Differ
- Context, Not Capability: Why a Smart AI Model Gives Poor Answers
- Discardable Software: When Code Is Cheap Enough to Throw Away
- Human in the Loop vs Human on the Loop: Supervising AI Agents
- The AI Audit Problem: When Making Work Is Cheaper Than Checking It
- AI Evaluation as Acceptance Test: How to Know It’s Good Enough
- Extensibility Is Control: Who Holds the Power in an AI Product
- RLHF Explained: How an AI Model Learns What People Prefer
- Pretraining Explained: How an AI Model Learns Before Anyone Teaches It
- Fine-Tuning Explained: How to Adapt a General AI Model to One Job
- Tokens Explained: The Unit AI Reads, Writes and Bills In
- Embeddings Explained: How a Machine Compares Meaning
- AI Hallucination Explained: Why Models State False Things
- Distillation Explained: How a Small Model Learns From a Large One
- Reasoning Models Explained: What Changes When AI Thinks First
- Tool Use Explained: How an AI Model Goes From Text to Action
- Capex and Depreciation Explained: Why Chip Lifetime Drives AI Profits
- Run-Rate Revenue Explained: What an AI Company’s Number Means
- Backlog Explained: Revenue That Is Signed but Not Yet Earned
- Switching Costs Explained: Why It Is Hard to Leave an AI Supplier







