A hardware argument from Positron AI’s Thomas Sohmers exposes a structural conflation at the center of almost every AI business debate: training spend and inference economics are different lines, and combining them hides more than it reveals.
What Happened
Speaking to Harry Stebbings on The Twenty Minute VC in an episode published on September 19, 2026, Thomas Sohmers — co-founder, CTO, and chairman of Positron AI, which builds hardware for AI inference — laid out three interlocking claims. On the hardware side, he stated that between 2014 and 2024, GPU compute improved by about 120× while memory bandwidth improved by only about 17×. Those are his own stated approximate figures. The ratio is not a rounding error — it is the arithmetic of a decade.
On the business side, Sohmers observed — citing reporting — that Anthropic’s API gross margin is reported to be around 80 points. That number is his relay of a reported figure, not an Anthropic disclosure, and it is not confirmed or endorsed here. He then made the broader argument, in his own words: “It’s more surprising to me how many people still today think that these are horribly unprofitable businesses and that the whole market’s going to zero,” and “Like if they stopped training, they’d be massively profitable overnight.” Those are his opinions, attributed to him here and not adopted.
He also disclosed his position before being asked: “those people are theoretically my customers and my margin is sort of based on their margin, but I care more about a healthy ecosystem long term.” That disclosure is reported here exactly as offered — as a transparency act, not a catch.
The key insight: The three claims — the hardware gap, the reported gross margin, and the unprofitability argument — are not separate observations. They are one argument about what happens when a single net number is used to describe two fundamentally different cost lines. The structure of that argument stands or falls independently of who is making it.

The Structural Read
The core of Sohmers’s argument is a point about accounting legibility. Training spend and inference operations appear on the same income statement, but they describe opposite conditions. Training is discretionary outlay directed at a capability that does not yet exist — it is an investment in a future state. Inference is the operation that serves paying demand today — it is the cost of the business that already exists. Combine them into a single net number and that number can no longer distinguish between a business that cannot cover the cost of serving its customers and a business that covers that cost comfortably while choosing to spend far ahead of its revenue on something else. Those are opposite conditions with opposite implications.
This is not a claim that any laboratory is profitable, would be profitable, or is solvent. No such claim is made here. Sohmers’s stronger version — that stopping training would make these companies massively profitable overnight — is his claim, reported as his. The structural point is narrower and more durable: a combined figure is not analytically inert, it is actively misleading, because it collapses the distinction between two opposite situations into a single sign.
The hardware numbers are where the cost curve finds its origin. On Sohmers’s figures, the arithmetic available on a single device grew roughly seven times faster over that decade than the bandwidth available to move data to it. The general property is well-established in systems design: when a workload is limited by moving data rather than by performing arithmetic, improvements in arithmetic do not relieve the limit. The binding constraint is the slower curve, and the cost per unit of work follows the bottleneck rather than the headline specification. That is why the 120× and 17× figures are not a technical aside — they are the reason the business discussion takes the shape it does.
Map of AI · Stack Layer Diagnosis
The Bottleneck Defines the Business Model
In the Map of AI stack, the hardware layer is not background infrastructure — it sets the cost floor for every layer above it. When memory bandwidth is the binding constraint, the economics of inference are determined not by how fast chips can calculate but by how slowly they can be fed. A gross margin figure at the application layer is therefore a downstream consequence of a bottleneck three layers below. Reading the gross margin without reading the bottleneck is reading the output without reading the mechanism.
On the reported gross margin figure: Sohmers, relaying reporting, places Anthropic’s API gross margin at around 80 points. That figure is not stated as fact here, and it is not confirmed by the company. What the figure illustrates — regardless of its precision — is a general property worth naming separately. A gross margin describes the unit economics of serving a customer against what that customer pays. It sits above every discretionary line below it — including training spend. It therefore says a great deal about whether serving demand is a productive operation and nothing at all, by itself, about the enterprise’s net result. Reading a high gross margin as a verdict on overall financial health is the same conflation as reading a combined net figure as a verdict on operations — it runs in the other direction but makes the same error.
Three Implications
IMPLICATION 1 · The P&L Is Not One Story
Anyone analyzing AI lab economics with a single combined net figure is not being conservative — they are being imprecise. The separation of training spend from inference operations is not a favor to the companies involved; it is the minimum resolution needed to ask whether the underlying business is working. That separation is available to any analyst who wants it. The fact that it is rarely performed in public commentary is itself a structural observation about how the debate is being conducted.
IMPLICATION 2 · Hardware Economics Propagate Upward Through the Stack
The memory bandwidth constraint is not a problem that software can fully route around. When the bottleneck is data movement rather than computation, model architecture choices, batching strategies, and inference optimization all operate within a ceiling set by the hardware layer. Companies building at the application layer — including the AI labs themselves — inherit that ceiling. A shift in memory architecture therefore propagates upward through every margin calculation above it, which is why a hardware argument and a business-model argument are the same argument told from different altitudes.
IMPLICATION 3 · Voluntary Disclosure Changes the Analytical Posture
Sohmers named his own exposure — the labs are theoretically his customers, and his margin tracks theirs — before being asked. An argument made by an interested party can be correct, and an argument made by a disinterested one can be wrong. The appropriate response to a voluntary disclosure is to assess the reasoning on its structure rather than to use the disclosure as a substitute for that assessment. The reasoning here has been assessed on its structure. The disclosure is reported because it is relevant, not because it resolves anything.
Get the structural read every week. Business Engineer covers AI economics, stack dynamics, and competitive positioning — the layer below the headline.
Subscribe to Business Engineer →The Bottom Line
Sohmers’s argument — a hardware observation, a relayed margin figure, and a challenge to the “horribly unprofitable” consensus — is structurally one claim: that combining two opposite cost lines into a single number produces a figure that cannot tell two opposite stories apart, and that most public commentary on AI lab economics has been reading that single figure as if it could. Whether his specific numbers are precisely right is a separate question from whether that structural observation is correct. The structural observation is correct. A bottleneck at the memory layer sets the cost floor; a gross margin at the application layer describes unit economics above every discretionary line; and a combined net figure, by design, destroys the distinction between the two. Analysts who want to assess these businesses need to recover that distinction before they reach for a verdict.
Not investment advice. Not a recommendation of any security, company, or course of action. This article is structural analysis only.
Sources: Thomas Sohmers on The Twenty Minute VC with Harry Stebbings, 19 September 2026 (YouTube). Gross margin figure is Sohmers relaying a report; not confirmed by Anthropic. Hardware figures are Sohmers’s own stated approximations. Analysis is original.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
This is not investment advice and not a recommendation. All quoted passages are Thomas Sohmers’s own words from the interview, checked against the caption track and an independent transcription. The figure of roughly 80 points of gross margin is Sohmers relaying a report — the word “reported” is his. It is not an Anthropic disclosure, it is not confirmed, no outlet is named above, and nothing above states it as fact. The claim that these companies would be massively profitable overnight if they stopped training is his argument, not a finding, and nothing above endorses it or says that any company is profitable, unprofitable, or burning cash. The 120x and 17x figures are his own stated approximations from the interview rather than measured benchmarks or vendor specifications. Sohmers volunteered, unprompted, that the laboratories are theoretically his customers and that his margin is in a sense based on theirs; that is reported above as a disclosure he offered, and nothing above characterises him as biased or conflicted. The arguments that training and inference are different lines of a profit and loss account, that a bottlenecked system follows its slower curve, and that gross margin sits above discretionary lines are general properties, not claims about any company’s actual results. Nothing is predicted.









