How a model becomes a dependable, economical service. Serve to the target. Keep the cache warm. Pay for busy, not idle.
Enter your email and the PDF opens. One step.
The book argues that in 2026 inference is where most of the money in AI is spent, and that most of the decisions that decide the bill are made by the people who build on models, not by those who make them.
It follows a single request from the moment it arrives to the moment its last token leaves, station by station, then steps back to the fleet, the economics and the failures. Each chapter can be read on its own by someone with a specific problem: an agent that has become too expensive, a chat product that feels slow, a fleet that is busy but unprofitable.
Time to the first token, how fast the rest arrive, how many requests meet their target, and what each accepted result costs.
Admission, batching, the cache, prefill and decode, faster decoding, smaller models, spreading across machines, and the cache beyond the accelerator.
When dedicated capacity beats paying per token, worked through on illustrative numbers you can replace with your own.
Gennaro Cuofano — creator of The Business Engineer & FourWeekMBA, read by 94,000+ founders, operators, and investors. Text and plates are original work. The book’s own advice: measure your own system before you believe any number in it.
119 pages, 44 plates, 36 chapters. Free — enter your email and it opens.
Get the next book and the weekly analysis, free
The Business Engineer is read by 94,000+ founders, operators and investors. One email, no cost, unsubscribe any time.
Subscribe free →Want the whole library?
The latest piece on this, Inference Engineering: The Tokenomics of AI, is for members. The opening is free to read. Right now membership is half price.
Read the opening → Join at 50% off →