Aleph Alpha’s Kolibri: 1M Context, Trained to 256K

Every figure and benchmark score here is Aleph Alpha’s own, from its blog post and Hugging Face card of 3 October 2026, run on its own harnesses. This publication verified none of it independently and did not read the full tech report.

Aleph Alpha released the weights of Kolibri, a 78B-parameter English-German model, on 3 October 2026 under Apache 2.0, and says it supports up to 1M tokens of context. Its own documents give the longest trained length as 256K, with 1M reached by a serving override. The benchmark scores are Aleph Alpha’s own, run on its own harnesses, and this publication verified none of them independently.

What Aleph Alpha Released

On 3 October 2026 Aleph Alpha released Kolibri, which its blog calls “an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active.” The weights are on Hugging Face under the Apache 2.0 license. The company’s model table gives 78.1 billion total and 3.46 billion active parameters, with 384 experts and 6 active per token.

Aleph Alpha says Kolibri is built for sovereign, mission-critical work in regulated areas including public administration, industrials and aerospace. It says an earlier 30.6 billion-parameter model, Kolibri Origin, had no public release.

Longest trained context for Kolibri Origin (65,536 tokens) and Kolibri (262,144), against the 1,048,576-token
Longest trained context for Kolibri Origin (65,536 tokens) and Kolibri (262,144), against the 1,048,576-token serving length Aleph Alpha documents as an override. All Aleph Alpha’s figures.

The Context Claim

The blog says Kolibri supports context lengths of up to 1M tokens. Its model table gives the longest trained length as 262,144 tokens, against 65,536 for Kolibri Origin, and the Hugging Face card says the model was trained on context windows up to 256k tokens.

Both documents say that serving contexts beyond 262,144 tokens needs a longer maximum model length and a configuration override. This publication’s own arithmetic: 1,048,576 is four times 262,144.

Aleph Alpha’s table gives LongBench Pro at 64.5 and AA-LCR at 68.3 for Kolibri. The benchmark table this publication read gives no score at 1M tokens.

How It Was Trained

Aleph Alpha says Kolibri was trained on 768 B200 GPUs in three stages: 20 trillion tokens of pre-training at a 16k sequence length over 21 days, 3.44 trillion tokens of mid-training at 64k, and 200 billion tokens of long-context adaptation at 256k, nearly 24 trillion tokens in all.

It says German accounts for about 4.3 trillion tokens, or 21.3% of the pre-training mix, against roughly 62% English and 14% code, and that the pipeline processed over 200 trillion tokens of raw data to reach the 20 trillion it trained on. It says pre-training finished on 11 September.

It says the 21-day run hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, handled automatically. This publication’s own arithmetic: 768 GPUs over 21 days is 387,072 GPU-hours, or about one interruption per 10,186.

The Benchmark Table

Aleph Alpha says its benchmarks “were run using our own harnesses and, where applicable, all models used the highest respective reasoning effort.” It compares Kolibri with Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B, and says Kolibri matches models with up to four times its active parameter count.

In its table Kolibri scores 96.9 on AIME 2025 against 91.7 for Nemotron 3 Super, 84.3 on GPQA Diamond against 83.4 for Qwen3.6, and 85.9 on LiveCodeBench v6 against 82.5. It trails Qwen3.6 on BFCL v4, 61.4 against 67.2, and on LongBench Pro, 64.5 against 70.8.

This publication’s own arithmetic from the 17 rows of that table: Kolibri has the highest score among Kolibri and the three named models on 10 rows and trails on 7. Aleph Alpha also says Kolibri sits on the Pareto frontier for quality against serving cost in English and German.

What Is Not Established

No third party has reproduced these scores in anything this publication read, and the harnesses are Aleph Alpha’s own. This publication did not read the full tech report, and the company gives no training cost, customer count or valuation.

All figures are Aleph Alpha’s own and none was independently verified. Nothing in this piece is investment advice.

This piece draws on Aleph Alpha’s blog post and Hugging Face model card of 3 October 2026. Every figure and benchmark score is the company’s own, run on its own harnesses, and none was independently verified. This publication did not read the full tech report. The 4x relation, the GPU-hour figures and the 10-versus-7 count are this publication’s own arithmetic. Nothing above predicts anything, and nothing here is investment advice.

Sources: aleph-alpha.com · huggingface.co · Aleph Alpha blog, ‘Kolibri Has Landed: A Sovereign Open-Weight Model’ (3 October 2026) · Hugging Face model card Aleph-Alpha

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA