When analytical tools are built around the tractable fraction of a genome — roughly one to two percent — the remaining 98 percent doesn’t get measured badly. It doesn’t get measured at all.
What follows concerns which fraction of the genome analytical tools have historically covered. It is not medical advice and makes no clinical claim, and it is not investment advice. It predicts nothing about discoveries, approvals or cures.
What Happened
On the Latent Space podcast, Eric Nguyen — co-founder and CEO of Radical Numerics, and co-developer of the HyenaDNA model line — offered a precise diagnosis of a structural problem in genomics: the analytical tools that the field built over decades have focused almost entirely on the protein-coding fraction of the human genome. That fraction, as confirmed by standard published references, sits at roughly one to two percent. Nguyen’s figure of “one and a half, two percent” is consistent with that established range, not a proprietary claim.
His second claim is different in kind and should be read differently. He characterised the non-coding remainder as the region where “many, if not most of the diseases” originate. That is an unquantified characterisation, not a measurement, and this article does not convert it into a percentage, a count, or a ranking. It is disclosed here as his framing — and returned to below — rather than treated as established fact.
Both disclosures about Nguyen matter: he has a commercial interest in long-context genomic models mattering (he leads a $50M-funded company building them) and an authorial interest in the account of why they were needed (he co-built the model he is describing). That is disclosure, not disqualification — and there is a specific reason it cuts less than it might, which the structural section below addresses.
The key insight: A short analytical window does not measure a long-range genomic relationship badly — it cannot represent one at all. Window size is the ceiling on what can be seen, not merely a parameter affecting precision. That distinction makes this an architecture story, not a new-biology story.

The Structural Read
This is a textbook case of what the Business Engineer Map of AI framework calls instrument-driven scope drift: an analytical estate gets built where measurement is tractable, and across enough time the map quietly becomes the territory. The questions a field asks drift toward the ones its instruments can answer. The tractable fraction of the genome — roughly one to two percent — attracted the tools. The rest did not disappear; it simply sat outside the window.
That is not a criticism of prior researchers. Coding regions were the rational place to begin: the relationship between sequence and protein is legible there, and considerably less legible outside it. You start where the reading is possible. The observation is about what a decades-long estate of short-context tools does to the shape of a field’s questions, not about any individual’s judgement. The same defect appears repeatedly across domains: a metric or instrument defined on a tractable subset silently narrows the scope of what gets investigated.
What changed is not biology but context length — and that distinction is the part worth actually understanding. Nguyen’s own description of the obstacle borrows from large language model practice: “context rot,” the tendency of models to degrade as input length grows. HyenaDNA, the model line he co-developed, is confirmed to carry context lengths up to one million tokens at single-nucleotide resolution. That architectural property matters because long-range relationships in a sequence are not detectable in principle from a short window — they are not measured imprecisely, they are simply invisible. Extending the window is not a refinement of the prior measurement; it is the precondition for a different class of measurement to exist at all.
Eric Nguyen — Latent Space Podcast
“If you’ve heard of the phrase context rot, you know, the longer the input you put into a language model, it starts to deteriorate. And so being able to pick up long range information and sort of patterns, motifs, grammar over long sequences was what we were trying to accomplish.”
Map of AI — Instrument Scope Layer
The Tractability Trap
Fields invest in infrastructure where outputs are interpretable. That investment compounds. Within a generation, the instruments define the research surface — not because the untractable territory is unimportant, but because it produces no signal the existing estate can read. Architecture changes that ceiling; it doesn’t improve the measurement, it removes the measurement’s prior constraint.
One specific reason Nguyen’s commercial interest cuts less than it might: the person who built the long-context architecture is also the person best positioned to identify what short-context tools could not see. The load-bearing factual number in his argument — the one to two percent coding share — is a standard published figure. The part that requires care is his characterisation of the non-coding remainder’s role in disease, which remains unquantified and is treated here as such.
Three Implications
IMPLICATION 1 — INSTRUMENT INVESTMENT IS SCOPE INVESTMENT
When a field funds tools, it is simultaneously choosing which portions of its subject are investigable. The 98 percent of the genome that sits outside the coding region did not go unmeasured by accident — it went unmeasured because the toolchain had no window long enough to hold a signal from it. Any new analytical infrastructure carries an implicit scope decision of the same kind.
IMPLICATION 2 — ARCHITECTURE BORROWS FROM AI, BUT THE TRANSFER IS LITERAL, NOT METAPHORICAL
The context-length constraint that Nguyen identifies is not an analogy to the language model problem — it is the same problem applied to a different sequence type. That means progress in long-context transformer and sub-quadratic architecture research has a direct, non-metaphorical application to genomics. The cross-domain transfer is technical, not rhetorical.
IMPLICATION 3 — FOUNDER-BUILT NARRATIVES REQUIRE STRUCTURAL DISCLOSURE, NOT STRUCTURAL DISMISSAL
Nguyen’s dual interest — commercial and authorial — is the normal condition for founder-led technical framing, not an exception. The appropriate response is layered disclosure: separate what is a standard published figure (the coding-region share) from what is an unquantified characterisation (the disease-origin framing), and treat each on its own evidentiary footing. Wholesale dismissal and wholesale acceptance are both analytical errors.
The Bottom Line
The Radical Numerics argument is not that prior genomics was wrong — it is that the analytical estate was built on the tractable one to two percent, and the window constraint was not a precision problem but a visibility problem: you cannot measure what your instrument cannot hold. Extending context length from hundreds to one million tokens is not a better measurement of the same thing; it is the removal of the constraint that made the rest of the sequence analytically invisible in the first place. Whether that capability translates into the outcomes Nguyen’s commercial framing suggests is a separate question, and an open one — but the architectural shift it represents is structural, not incremental.
Sources: Eric Nguyen on the Latent Space Podcast (YouTube); standard references on protein-coding DNA share of the human genome are consistent with the ~1–2% figure cited.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Eric Nguyen is co-founder and chief executive of Radical Numerics, which builds models of the kind discussed above and has raised $50 million, and he co-developed the model line he is describing. Both interests — commercial and authorial — are disclosed here rather than treated as disqualifying. His statement that many or most disease-causing variants sit in non-coding regions is a characterisation rather than a measurement. It is unquantified, and no figure is put on it above. His figure of about one and a half to two per cent for coding regions is consistent with the standard reference figure of roughly 1 to 2 per cent for protein-coding DNA. The share shown above is a share of the genome, not a share of disease. The two are not the same quantity and are not treated as such. No sequence design, synthesis, organism, biological agent or method of any kind — computational or otherwise — is described or named above. The article addresses only which fraction of the genome analytical tools have historically covered. Nothing above is medical advice and nothing above makes a clinical claim; no disease is named. Model accuracy and benchmark figures, any named disease, gene, variant or patient population, any clinical or regulatory status, any revenue, customer or partnership, any count of variants and any competitor comparison are not established and do not appear. Nothing above predicts discoveries, approvals, cures or the direction of the field, and nothing above is investment advice.









