Voltropy Vast-10M Wins at 1M Tokens, Not at 100K

Voltropy’s Vast-10M-Flash passes Claude 5.1 Fable at one million tokens — but its own technical report shows Fable leads by eight points at 100,000. The story is decay, not dominance.

Every score below is Voltropy’s own, from its own run of a third-party benchmark. This publication has run nothing and reproduced nothing. Vast-10M-Flash does not lead at short context. At 100,000 tokens Claude 5.1 Fable scores 60.98 and GPT-6 Astra 55.56, against Flash at 52.66. It passes Fable at one million tokens because Fable falls further, not because Flash improves. Nothing here is investment advice.

What Happened

Voltropy launched Vast-10M and its Voltropy Sparse Attention (VSA) library on 30 September 2026, in early access. The headline claim, carried by PR Newswire, was that Vast-10M-Flash beats Anthropic’s Claude 5.1 Fable on the BEAM long-context benchmark. Every figure in this article comes from Voltropy’s own technical report. Voltropy ran the evaluation. This publication has run no benchmark and reproduced nothing.

The claim is true — at one context length. At 100,000 tokens, Claude 5.1 Fable scores 60.98 on BEAM. GPT-6 Astra scores 55.56. Vast-10M-Flash scores 52.66, behind both. At 500,000 tokens the order holds: Fable 57.26, Astra 56.15, Flash 50.22.

At one million tokens the ranking inverts. Astra reaches 49.75. Flash reaches 48.73. Fable falls to 44.03. Flash passes Fable without ever getting better than it at shorter lengths. The crossover is real. The conquest framing is not.

The key insight: Fable drops 16.95 points between 100,000 and one million tokens — a fall of 27.8 percent. Flash drops 3.93 points, or 7.5 percent. The crossover happens because Fable decays faster, not because Flash accelerates. Nothing about Fable changed. The test got longer.

Every figure is Voltropy's own, produced by its own run of a third-party benchmark. Nothing here has been repr
Every figure is Voltropy’s own, produced by its own run of a third-party benchmark. Nothing here has been reproduced independently. DeepSeek V4.1 Flash is the one model that does not fall at all.

The Structural Read

A context window is usually quoted as a capacity. Ten million tokens. One million tokens. A number on a spec sheet. What Table 1 in Voltropy’s report actually measures is something different: decay. How much of a model’s reasoning survives as you fill the window.

Those are different properties. They rank differently. And the industry has been conflating them since context windows became a marketing variable.

Voltropy’s real product claim belongs on the decay axis. Vast-10M-Flash scores 40.20 at ten million tokens, retaining 82.49 percent of its one-million-token score. Its base model, DeepSeek V4.0 Flash, scored 39.69 at one million tokens — below what Flash achieves at ten times that length. That is the technically interesting result. It is also, notably, evaluated on 200 questions, the smallest of the four sample sets.

There is a further complication the report does not fully resolve. An unmodified model Voltropy did not convert — DeepSeek V4.1 Flash — scores 46.00 at one million tokens. That is above both Vast-10M-Pro at 43.76 and Vast-10M-Medium at 45.67. The report argues that Vast-10M-Flash still beats V4.1 Flash at every shared length, and Table 1 supports that. But the pro and medium variants do not.

The report’s methodology is more transparent than most vendor benchmarks. It pins the evaluation to a specific BEAM commit: git revision 3e12035532eb85768f1a7cd779832b650c4b2ef9. It reports full runs at three of four tiers — all 400 questions at 100,000 tokens, all 700 at 500,000, all 700 at one million. It names the judge: GPT-4.1-mini at temperature zero, using BEAM’s official scoring formula.

BEAM itself is a third-party benchmark from Tavakoli and colleagues, arXiv 2510.27246, accepted at ICLR 2026. A third-party benchmark evaluated by the vendor. Both halves of that matter.

The report also names its own central confound: retrofitting an attention mechanism requires retraining, and a score gain could be an artefact of that retraining acting as a fine-tune rather than evidence that VSA improves anything. A clean test would fine-tune the original checkpoint on the same corpora and compare. Voltropy did not run that test. It asserts that VSA required little retraining and that the ablation was therefore unnecessary.

That is an assertion, not a measurement. Naming the confound is more than most vendors do. The check it describes is still missing.

One more absence is worth stating plainly. The report contains no efficiency figure. No parameter count. No latency. No throughput. No memory footprint. No cost per token. For a new attention mechanism whose core claim is performance at scale, that is a conspicuous gap. The reason for it is not stated in the document.

Voltropy Technical Report — Sep 30, 2026

“Retrofitting a new attention mechanism inherently involves some degree of retraining… We believe this was not a significant factor as VSA did not require extensive retraining.”

Three Implications

BUYERS

A context window quoted in tokens tells you capacity. It does not tell you what score a model still holds when that capacity is used. The useful procurement question is now decay rate at the length you actually run — not the headline window size. That number is rarely on a model card.

ANTHROPIC AND OPENAI

Fable’s 27.8 percent drop from 100,000 to one million tokens, compared with Flash’s 7.5 percent, is now a published number with a specific benchmark citation. Whether that gap reflects architecture, training data, or evaluation methodology, Anthropic has a documented decay profile to respond to. Astra’s 10.5 percent drop places OpenAI between the two.

BENCHMARK DESIGN

The BEAM structure — different numbers of questions at different tiers — means scores across tiers are not directly comparable. They measure related but distinct tasks. The field needs a benchmark that holds task construction constant as context grows. Until it exists, crossover claims like this one will always carry a methodological asterisk.

Business Engineer Framework

Product Overhang Doctrine — Map of AI

Decay rate has been an invisible specification accumulating in the background while the industry marketed window size. Voltropy’s Table 1 is the moment it surfaces. The Map of AI framework maps exactly where infrastructure plays like this sit in the nine-layer AI stack — and which layer captures the value when a hidden spec becomes visible.

Explore the Map of AI →

The Bottom Line

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Every score above is Voltropy’s own, taken from Table 1 of its Vast-10M Technical Report, downloaded as a PDF from voltropy.com and read on 30 September 2026. This publication has run no benchmark and has reproduced none of these results. BEAM is a third-party benchmark, from Tavakoli, Salemi, Ye, Abdalla, Zamani and Mitchell, arXiv 2510.27246, accepted at ICLR 2026. The evaluation reported here was run by Voltropy, not by the benchmark’s authors and not independently. Nothing above should be read as Voltropy beating Anthropic’s Fable 5.1 or OpenAI’s GPT-6 Astra in general. On Voltropy’s own table, Vast-10M-Flash is behind both at 100,000 tokens and behind both at 500,000 tokens. It passes Fable at one million tokens because Fable’s score falls further over that range, not because Flash improves. On the ablation: the report itself raises the possibility that a gain from retrofitting an attention mechanism could be an artefact of the retraining required, describes the fine-tuning control that would settle it, and states that it did not run that control because VSA did not require extensive retraining. That is the company’s assertion rather than a measurement, and nothing above suggests the point was concealed — Voltropy raised it. The ten-million-token results rest on 200 questions, the smallest of the four tiers. The report contains no parameter count, no latency, no throughput, no memory footprint and no cost figure, and gives no reason for their absence; nothing above speculates about one. The launch was also distributed over PR Newswire, which is paid distribution rather than independent reporting. Nothing above predicts anything, and nothing here is investment advice.

Sources: voltropy.com · openrouter.ai · fourweekmba.com · voltropy.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA