Mathematician Daniel Litt draws a sharp line between generating correct artifacts and producing the thing mathematics is actually for.
On the a16z Podcast episode “Can AI Learn Mathematical Intuition?”, Litt — a working mathematician at the University of Toronto — makes an argument that cuts through a lot of AI-and-science optimism. His position is not that AI cannot produce correct mathematics. His position is that correctness is not the point.
The real goal, in his account, is understanding — the kind a human mathematician develops through curiosity, failure, and the hard work of internalizing why something is true. If that understanding is latent somewhere in a model’s weights, Litt finds that answer unsatisfying. Not wrong. Unsatisfying. That distinction matters enormously for how we think about AI in any knowledge-intensive domain.
The key insight: When the measurable proxy — papers, proofs, correct outputs — diverges from the real goal — understanding, insight, the capacity to ask the next right question — optimizing the proxy is not progress. It is a very sophisticated form of slop.
The Structural Read
AI systems are extraordinarily good at the output layer of knowledge work. They can generate proofs, papers, code, analysis — artifacts that look like the product of expertise. What Litt’s argument surfaces is that in mathematics, and arguably in most deep knowledge work, the artifact is a byproduct. The real deliverable is the mental model that produced it.
This creates a structural problem as AI scales into research. The moment you can generate volumes of correct-ish artifacts at low cost, the incentive structure warps. Quantity becomes easy to demonstrate; quality of insight becomes harder to distinguish from it. Litt’s complaint is not about hallucinations. It is about a subtler gap — between a system that can close a proof and one that knows which theorem is worth proving.
That gap is not unique to mathematics. It shows up in any domain where the easy-to-measure output — the report, the recommendation, the generated strategy — diverges from the thing that actually matters: the judgment that selected it.
ARTIFACTS ARE NOT UNDERSTANDING
Litt’s core claim, as a working mathematician, is that mathematics exists to build understanding — not to produce papers. Even if a model encodes some form of understanding in its weights, that is, in his view, a different and lesser thing than the understanding a curious human develops and can act on.
THE PROXY PROBLEM SCALES
When AI can generate high volumes of correct-looking output cheaply, organizations face a harder problem: the proxy metric (output produced) becomes easier to hit while the real metric (insight quality, theorem selection, judgment) becomes harder to audit. The gap between the two is where value quietly erodes.
CURIOSITY IS THE ACTUAL CAPABILITY
Litt locates the human advantage not in correctness but in motivation and direction — the capacity to be genuinely curious, to choose the right problem, and to build a personal understanding that evolves. That is the capability AI in research has not replaced; it is also the one that is hardest to benchmark.
The Bottom Line
Daniel Litt is not making a doomsday argument or a dismissal. He is making a precise one: correctness and understanding are not the same thing, and mathematics — like most serious knowledge work — is in the business of the latter. The fact that AI can produce the former at scale is genuinely impressive. It is also, in Litt’s framing, beside the point. That distinction should sit at the center of any honest conversation about what AI is actually replacing and what it is not.
Clip via the a16z Podcast, “Can AI Learn Mathematical Intuition?” — Daniel Litt (mathematician, University of Toronto) with host Lisha Li (a16z).







