Illumina introduced SpliceAI2 on 8 October 2026, a genomic AI model that predicts how DNA variants change RNA splicing, the step in which RNA transcripts are assembled before protein production. Illumina aims it at rare disease research, hereditary cancer testing research and drug discovery.
In a preprint from Illumina’s BioInsight AI Lab, SpliceAI2 identified 17% more disease-relevant splice variants than other models in analyses of Genomics England cohorts. The code and trained models are on GitHub for non-commercial organisations only; commercial customers reach the model through Illumina’s BioInsight applications.
What SpliceAI2 Predicts
Splicing is how a cell cuts and joins RNA before it makes protein, and Illumina’s article says over 95% of human genes undergo alternative splicing. The original SpliceAI, which the same lab released in 2019, answers one question, the article says: does a cell splice at a given location or not.
SpliceAI2 is built to answer three, the article says: which positions are used as splice sites and how often, which splice sites connect to one another, and which full-length RNA transcripts result.
Illumina says all that is needed to run SpliceAI2 is a DNA sequence. The stated aim is to spare researchers the cost of generating long-read RNA data, and the need for RNA from tissues that are hard to obtain.

Business Pill · OPEN WEIGHTS
A one-minute explainer of open weights: what it means when a model’s trained parameters are published, and what licences can still restrict. It teaches the general idea only and says nothing about any company in this story.
The key insight: As we read it, SpliceAI2 reaches users by two routes: the code and trained models on GitHub, which Illumina’s article lists and whose terms of use restrict to non-commercial organizations such as universities, academic research institutions and government entities; and Illumina’s BioInsight applications, which the release names as the way customers access SpliceAI2.
The 17% and the Other Figures
The headline figure comes from the Genomics England 100,000 Genomes Project. Illumina’s team examined 7,504 individuals with extensive phenotype information and asked whether variants SpliceAI2 flags fall more often in genes already linked to each person’s condition.
At matched confidence thresholds, SpliceAI2 identified 17% more disease-associated variants than any other tested splicing model, the article says. Against the original SpliceAI, it found 33% more disease-relevant splice variants at what Illumina calls a 2X confidence interval, and 66% more at 4X.
On matched genome and RNA sequencing data from the NIH’s Genotype-Tissue Expression Portal, the release says SpliceAI2 improved quantification of splice site usage by 34% compared with the next best model.
The comparisons covered the original SpliceAI, Pangolin and Google DeepMind’s AlphaGenome. Illumina says the AlphaGenome comparisons were run independently by academic collaborators at the University of Oxford.
Roughly 50% of the cryptic splice variants SpliceAI2 identified sat deep within introns, the article says. The preprint puts them more than 50 base pairs into introns, where exome sequencing would typically miss them.
Where the Gains Come From
Illumina credits the training data. SpliceAI2’s training set was more than 100 times larger than the original’s: 314,745 RNA sequencing samples across ten species, covering more than 46 million observed splice junctions after filtering, the article says.
For full transcripts, the lab added 330 long-read samples from the public ENCODE project. With them, the article says, the model reconstructed the single most common transcript 82% of the time for genes it had never seen in training, against 78% when trained on short-read data alone.
The model also takes in the measured activity levels of 147 splicing regulators. Across nearly 15 million measurements spanning 48 human tissues it captured tissue-specific splicing patterns, the article says, but it was less successful at predicting how the impact of a given variant changes between tissues.
Population Checks and a Three-Model Suite
Illumina also ran SpliceAI2 on a combined cohort of 627,000 genomes and 36,764 proteomes drawn from gnomAD, TOPMed and UK Biobank. The article says the highest-scoring variants were strongly depleted across those genomes, and carriers of higher-scoring variants had lower plasma protein levels.
In the rare disease data, the preprint says protein-truncating variants accounted for 50% of the excess genetic burden, cryptic splice variants identified by SpliceAI2 added 15%, and the remainder came from variants identified by Illumina’s PrimateAI-3D and PromoterAI.
The release says SpliceAI2, PromoterAI and PrimateAI-3D together “enable researchers to identify up to twice as many variants with predicted biological impact.” Kyle Farh, vice president of Illumina’s BioInsight AI Lab, said: “Illumina is advancing AI to systematically shrink the portion of the genome that remains uninterpretable.”

Open to Universities, Built Into Illumina Software
Source code, trained models and precomputed predictions for every possible single-nucleotide variant within human gene bodies are on GitHub, Illumina’s article says. The terms of use in that repository set the limits.
Under them, SpliceAI2 “is only made available for non-commercial use by, or on behalf of, non-commercial organizations such as universities, academic research institutions, and government entities.” The terms exclude for-profit companies, contract research organizations and laboratories acting for commercial purposes.
For customers, the release says SpliceAI2 is reached through Illumina’s BioInsight applications, naming DRAGEN Annotation and Emedgene; the article adds Illumina Connected Insights. BioInsight is Illumina’s data, software, informatics and AI business, the release says.
The first SpliceAI has been cited in more than 3,400 publications and is incorporated into ClinGen’s guidelines for splice variant interpretation, the release says.
The Structural Read
Illumina’s article attributes the gains largely to training data: a set more than 100 times larger than the original SpliceAI’s, drawn from 314,745 RNA sequencing samples, plus 330 long-read samples for full transcripts. As we read it, the model’s edge rests on assembling that data, not on access to the code.
The 17% headline figure and the 15% share of the excess genetic burden come from the same Genomics England analysis, as reported in Illumina’s article and preprint. Both describe research findings; the release frames the use as rare disease research, hereditary cancer testing research and drug discovery.
The release presents SpliceAI2 as one of three Illumina models alongside PromoterAI and PrimateAI-3D, and says together they enable researchers to identify up to twice as many variants with predicted biological impact. As we read it, the three models are offered as one annotation stack inside Illumina’s own applications.
Rami Mehio, senior vice president and general manager of BioInsight at Illumina, in the release (8 October 2026)
“As researchers work to elucidate the effect of mutations, we are uniquely positioned to unite genomic data and scientific expertise at scale, delivering the AI tools that can advance discovery and human health.”
Three Implications
ACADEMIC AND GOVERNMENT LABS The terms of use make the code, trained models and precomputed predictions available to non-commercial organisations such as universities, academic research institutions and government entities.
COMPANIES AND COMMERCIAL LABS The same terms exclude for-profit companies, contract research organizations and laboratories acting for commercial purposes; the release names DRAGEN Annotation and Emedgene as the applications through which customers reach SpliceAI2.
OTHER SPLICING-MODEL BUILDERS Illumina’s comparisons named the original SpliceAI, Pangolin and Google DeepMind’s AlphaGenome, with the AlphaGenome runs done by collaborators at the University of Oxford, the article says.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Open vs Closed Meta-Framework.
The framework puts the pattern this way: “Close the SCARCE layer (keep proprietary) + Open the ABUNDANT layer (commoditize)”. It adds that “identifying the right layers is only half the battle.”
As we read it, SpliceAI2’s two routes follow those lines: Illumina’s article lists code and trained models on GitHub, which the terms of use restrict to non-commercial organizations, while the release names Illumina’s BioInsight applications as the customer route. The article credits the gains largely to training data drawn from 314,745 RNA sequencing samples.
What Is Not Established
Every performance figure here is Illumina’s, from its release, its article and a preprint by Illumina’s BioInsight AI Lab. We have not run the model.
Neither the release nor the article states a price for access through the BioInsight applications. Illumina’s preprint also names a limit: “Further progress is needed to predict how splice-altering variant effects vary across tissues, cell types, and developmental contexts.”
The Bottom Line
Illumina’s SpliceAI2, released on 8 October 2026, predicts splice sites, splice junctions and full transcripts from a DNA sequence alone. In Illumina’s own analysis of Genomics England cohorts, it identified 17% more disease-relevant splice variants than the other tested models. Universities and other non-commercial organisations can use the code and trained models from GitHub; commercial users reach it through Illumina’s BioInsight software.
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. We read Illumina’s release and its article on SpliceAI2 in full on 10 October 2026, along with the abstract, rare-disease results and discussion of the BioInsight AI Lab preprint and the terms of use in the GitHub repository; every performance figure here is Illumina’s. We did not run the model. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: Illumina: Illumina releases SpliceAI2 to help advance rare disease research (PR Newswire, 8 Oct 2026) · Illumina: Introducing SpliceAI2 (article, 8 Oct 2026) · Illumina BioInsight AI Lab: A unified framework for quantitative splicing and transcript prediction (preprint) · Illumina: SpliceAI2 repository and Model Terms of Use (GitHub)









