TypeSafe AI’s Diogo Almeida Argues ChatGPT and Claude Code Belong to the Same Era

A co-author of ChatGPT and RLHF argued at AI Engineer that the assistant era and the coding-agent era share a defining structural property — and that the field may be drawing its era boundaries in the wrong place.

Context at a Glance

31 Jul 2026

Talk uploaded to AI Engineer

RLCD

TypeSafe AI’s described training method for Jev

GPT-4

One of Almeida’s co-authorship credits

Jev

TypeSafe AI’s described “System One Model”

What Happened

Speaking at the AI Engineer conference in a talk uploaded on July 31, 2026, Diogo Almeida — a co-author on GPT-4, ChatGPT, and the RLHF and InstructGPT papers, and someone who said on stage that the team he was part of basically invented post-training as a concept — opened by collapsing a distinction the industry largely treats as settled. He is now co-founder and chief executive of TypeSafe AI, and he was previously a researcher at Google Brain and OpenAI.

The argument was structural, not capability-focused. Almeida did not claim RLHF is over, dead, or done — no such statement appears in the talk. His claim was about era boundaries: specifically, that the industry is drawing them around capability levels when it should be drawing them around output types and the position of the human in the transaction.

TypeSafe AI is separately reported to be building a system called Jev, which the company describes as a “System One Model” trained with RLCD — reinforcement learning for calibrated decisions — that returns structured decisions with calibrated confidence rather than generated text. That is the company’s own characterisation of its own product. Secondary coverage, though not this talk, also reports Almeida as arguing the field will eventually look back on the ChatGPT era as “a weird detour.”

Diogo Almeida — AI Engineer, July 31 2026

“I’m talking about what’s next after RLHF. More accurately, I think this should be called what’s next after the ChatGPT era that I think we’re all in. And my hint for you guys is it is not the Claude Code era. I will justify this later on, but I actually believe them to be part of the same era.”

The key insight: A person revising the significance of his own work is making a categorically different kind of claim from an outsider forecasting a rival’s decline. Almeida is not predicting that someone else’s approach will fade — he is arguing that the thing he helped build occupies a different place in the sequence than the field currently assumes. That is a structural claim, and it deserves to be weighed on its structure.

The Structural Read

The Map of AI framework asks where value is created and captured across the stack — and more precisely, what property of a layer determines when that layer has been superseded. Almeida’s claim provides a candidate answer that is worth taking seriously on its own terms, independently of whether it is correct.

A chat assistant and a coding agent differ enormously in scope, autonomy, and practical usefulness. Nobody sensible disputes that a coding agent does more. What they share is the shape of the transaction. The system produces an artefact — prose, a plan, a diff — and a human being reads that artefact and decides whether to accept it. Every advance that makes the artefact longer, better, or more ambitious is an advance within the era rather than a transition out of it, because the position of the human has not moved. The human is still the adjudicator. The system is still producing something for adjudication.

On this reading, the era ends not when capability increases but when the output type changes: specifically, when a system returns something that can be acted on without being checked. That is a different bar. It requires not fluency, not capability, and not being usually right — it requires a reliable signal about the system’s own reliability. A system that is right most of the time but cannot tell you when it is wrong still requires a human reviewer for every single output, because the reviewer has no way of knowing in advance which outputs are the exceptions. Accuracy reduces the cost of review. It does not remove the need for it. Only a reliable signal about its own reliability could do that, and that signal is a separate property from the answer itself.

Map of AI — Era Boundary Thesis

The Adjudicator Position

If eras are defined by output type rather than capability level, then the relevant boundary is not “how good is the artefact?” but “does the human still have to check it?” Longer, better artefacts are within-era advances. The era transition requires a shift in who bears the decision — and that shift requires calibrated confidence, not just higher accuracy. This is a general property of delegation under uncertainty; nothing here asserts any existing system has or lacks it.

The honest note on Almeida’s position is that a founder whose product thesis matches his public diagnosis of the field is the ordinary case, not a disqualification. People build what they believe is missing, and the belief usually precedes the company. That is also a reason to weigh the argument on its structure rather than on the strength of the credentials attached to it — which is the better way to evaluate any thesis of this kind, regardless of who is making it.

Three Implications

IMPLICATION 1 — HOW LAYERS ARE EVALUATED

If the era boundary is output type rather than capability level, the correct question for evaluating any AI layer is not “does it perform better?” but “has the shape of the transaction changed?” Products that make artefacts substantially better are still competing within the same paradigm. Products that shift where the human sits in the loop are competing for a different position entirely — and the competitive moat for the latter is correspondingly wider if it can be established.

IMPLICATION 2 — CALIBRATION AS A SEPARATE ENGINEERING PROBLEM

The argument implies that calibrated confidence — knowing when an output can be trusted — is a distinct engineering target from accuracy or fluency, and that current training paradigms do not automatically produce it as a byproduct of higher capability. Whether that is true is an empirical question this piece does not answer. But if it is true, it reframes a significant portion of the current research agenda: gains in benchmark performance do not close the delegation gap if they do not also improve the reliability of the system’s self-signal.

IMPLICATION 3 — THE CREDENTIAL CUTS BOTH WAYS

The fact that a co-author of the foundational systems is the one proposing this reframing gives the argument a different weight than an outsider’s prediction — but it also makes it harder to separate intellectual conviction from strategic positioning. The right response is not to discount the argument because of who is making it, nor to accept it because of who is making it. It is to examine whether the structural property it identifies — the adjudicator position — is real, and whether it tracks the distinction it claims to track.

Business Engineer Framework

The Map of AI

Almeida’s era-boundary argument is a question about where a layer sits and when it has been superseded — exactly the problem the Map of AI is built to answer. The framework maps 200+ companies across 9 layers of the AI stack, and the most important question it asks for each layer is not “is it getting better?” but “what structural position does it hold and what would displace it?” Understanding the adjudicator position is a Map of AI question.

Explore the Map of AI →

The Bottom Line

Whether Almeida is right about where the era boundary sits is an open question this piece does not resolve. What the argument does, usefully, is name the property that would have to change: not capability, not fluency, not benchmark performance, but the position of the human — from adjudicator of every output to trustee of a system that knows when it can be trusted. That is a precise target, it is a hard one, and the fact that a co-author of the foundational systems thinks it has not yet been reached is worth sitting with, regardless of what comes next.

This article is analysis of a public talk and is not investment advice. No benchmark, result, accuracy figure, valuation, raise, headcount, customer, launch date, roadmap, or market size has been stated or implied. No existing system is named as calibrated or uncalibrated. TypeSafe AI’s descriptions of Jev, “System One Model,” and RLCD are the company’s own characterisations of its own product and are attributed as such throughout.


Sources: Diogo Almeida — AI Engineer talk, uploaded July 31 2026 (YouTube); analysis by Business Engineer / FourWeekMBA editorial team, September 19 2026.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This is analysis of a public talk. It is not investment advice and not a recommendation. The quoted passage is Diogo Almeida’s own words from his AI Engineer talk, checked against the caption track and an independent transcription. He did not say that RLHF is over, dead or done, and nothing above should be read as reporting that he did; his stated claim is about what comes after, and about two supposed eras being one. The phrase “a weird detour” appears in secondary coverage of his views rather than in this talk. Jev, the “System One Model” label and RLCD are TypeSafe AI’s own descriptions of its own product, relayed here as such; nothing above says the approach works, performs, outperforms, is reliable, or is in fact calibrated, and no benchmark or result appears. No model is named as calibrated or uncalibrated. Nothing above says Almeida is right or wrong, accuses him of talking his book, or endorses his product, and nothing above claims what any other company intends. Nothing is predicted.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA