When the binding constraint is the megawatt, a fivefold efficiency gain buys five times the compute — not a fifth of the power bill. The denominator is the whole argument.
What Happened
On a recent episode of Harry Stebbings’ 20VC, Thomas Sohmers — founder and CTO of Positron AI, which sells inference accelerators that compete with Nvidia equipment — was asked whether energy is the binding constraint on AI buildout. His answer was more candid than most efficiency pitches tend to be. “On our base case,” Sohmers said, “if we can turn what you would have spent 500 megawatt with Nvidia equipment and do that in 100 megawatt, I don’t think that’s actually going to mean that you’re only going to build a 100 megawatt facility — you’re still going to build the maximum amount of compute that you can, you’re just getting more tokens, more intelligence per joule.” That 500-to-100 figure is a conditional base-case hypothetical framed explicitly as “if we can” — it is not a measurement and not a shipped benchmark. Sohmers has a direct commercial interest in the efficiency argument being taken seriously; that is disclosure, not disqualification, and it matters for reasons returned to below.
Positron’s own published materials lead primarily with dollar-denominated comparisons: 24.8 times revenue per total-cost-of-ownership dollar, and 26 times tokens per dollar at a matched generation speed of 170 tokens per second, alongside a throughput figure of 400 tokens per second per user against Nvidia Blackwell GB300 NVL72. Those are the company’s own claims and have not been independently verified here. The megawatt framing is the register Sohmers reaches for in conversation rather than in the published headline numbers — a difference in denominator that is worth examining on its own terms, with no motive imputed to anyone.
The reason the interview answer is worth slowing down for is structural. Sohmers is not making a consumption claim dressed up as an efficiency claim. He is stating plainly that the efficiency gain will be spent as more compute rather than banked as less power — and in doing so, he is giving up the environmental sales line that was available to him. Arguments that cost their maker something are usually the ones that reward closer reading.
The key insight: Efficiency is a ratio. A ratio improves either by shrinking the numerator or by growing the denominator. Which of those actually happens is not a property of the technology — it is decided by which term is constrained. When the megawatt is the binding constraint and compute is the free variable, a fivefold efficiency gain produces five times the compute at the same draw. The gain is spent, not banked.

The Structural Read
The Map of AI framework treats the infrastructure stack as a set of discrete layers — each with its own scarce input, its own buyer, and its own way of pricing a gain. Positron’s efficiency claim lands differently depending on which layer you are standing on, because each layer names a different denominator as the thing it is short of.
An operator whose scarce inputs are capital and site capacity reads a fivefold efficiency gain as an unambiguous win: more output per dollar, more output per megawatt-hour, better unit economics on the tokens that generate revenue. The gain is real and the framing is accurate. A siting authority or grid planner whose scarce input is the megawatt itself reads the identical gain differently — the draw is unchanged by construction, because filling the power envelope is the assumption that produced the number in the first place. These are not competing claims and neither party is being imprecise. They are the same ratio read from two different sides of the constraint.
This is also why the denominator shift between Positron’s published materials and Sohmers’ interview framing is analytically interesting rather than suspicious. A dollar ratio is a cost claim aimed at a CFO or a procurement desk. A tokens-per-joule ratio is a thermodynamic claim aimed at an engineer. A megawatt count is a siting claim aimed at a land and power team. They are not interchangeable, and a buyer, a regulator, and a grid planner are each asking for a different one of the three. The confusion in most public argument about AI and energy comes from treating a throughput claim as though it were a consumption claim — reading the operator’s win as though it were the grid planner’s win.
Thomas Sohmers — Positron AI, via 20VC
“You’re still going to build the maximum amount of compute that you can, you’re just getting more tokens, more intelligence per joule.”
Map of AI — Infrastructure Layer
The Constrained-Optimisation Point
In constrained-optimisation terms, efficiency gains are always consumed by whichever constraint binds. When power is the ceiling, better silicon raises the compute floor, not the power ceiling. The gain is real — but where it lands depends entirely on what you are short of, not on what the silicon is capable of. Any regulatory cap expressed in megawatts binds on power draw directly; a throughput improvement converts that cap into more tokens, not fewer megawatts.
Three Implications
FOR OPERATORS: The efficiency gain is real and it compounds
If the site power envelope is fixed — which it usually is, given interconnection timelines — then more compute per megawatt is a direct improvement in revenue per square foot of data center. The operator wins on both cost (the dollar-denominated multipliers Positron leads with) and on capacity (more tokens from the same grid allocation). The claim is aimed squarely at this buyer, and the framing is accurate for it.
FOR BUYERS AND PROCUREMENT TEAMS: The denominator you ask for determines the answer you get
A dollar-per-token comparison, a watt-per-token comparison, and a peak-throughput comparison are three different questions with three different answers. Positron’s published multipliers are the company’s own figures and are not independently verified here — but the more general point holds for any vendor comparison: require the denominator that matches your actual scarce input before treating efficiency ratios as comparable across competing claims.
FOR POLICY AND GRID PLANNING: A megawatt cap converts into more tokens, not fewer megawatts
Any regulatory instrument expressed in megawatts — a cap, a tariff trigger, a zoning threshold — is not softened by silicon efficiency improvements when operators fill their power envelopes. Sohmers states this plainly. The policy implication is not that efficiency silicon is bad for grid planning; it is that a megawatt-denominated rule and a tokens-per-joule improvement operate on different variables entirely and do not automatically trade off against each other.
The Bottom Line
Sohmers’ interview answer is worth reading carefully not because the conditional hypothetical is a benchmark — it is explicitly not — but because he states the constrained-optimisation logic out loud rather than letting an audience assume that efficiency means reduction. A dollar ratio, a thermodynamic ratio, and a siting ratio are three different questions; confusing them is where most public argument about AI and energy goes wrong. The denominator is the whole argument, and whoever gets to name it controls what the number means.
Sources: Thomas Sohmers on 20VC with Harry Stebbings (YouTube); Positron AI published materials (company’s own claims, not independently verified). This article is not investment advice.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
The 500-megawatt-to-100-megawatt figure is Thomas Sohmers’s conditional base-case hypothetical, not a measurement or a shipped benchmark. His own framing was “on our base case, if we can”, and nothing above treats it as a benchmark result. Sohmers is founder and chief technology officer of Positron AI, which sells inference accelerators competing with Nvidia equipment, so he has a direct commercial interest in the efficiency argument. That is disclosure rather than disqualification — the argument stands or falls on whether the constrained-optimisation point holds. Positron’s published multipliers — 24.8× revenue per total-cost-of-ownership dollar, 26× tokens per dollar at matched speed, and 400 against 170 tokens per second per user versus Nvidia Blackwell GB300 NVL72 — are the company’s own claims and are not independently verified here. The approximate 2.35× throughput ratio is arithmetic performed by the author from those two published figures. No concealment, burial, omission or downplaying is alleged against Positron or anyone else, no claim is made that a watt-denominated comparison does not exist, and no motive is imputed to any company or person. The difference in denominator is described, not characterised. Independent benchmarks of Positron silicon, any Nvidia response or figure, any operator’s actual power draw, Positron’s revenue, customers, shipments, funding and valuation, any grid operator’s position, interconnection queue lengths and electricity prices are not established and do not appear. Nothing above predicts datacentre buildout, power prices, chip adoption, emissions or regulation, and nothing above is investment advice.









