TeleOCR at 1.2B Parameters Tops a Document-Parsing Benchmark

A purpose-built 1.2B-parameter model reports outscoring systems more than 200 times its size on one narrow, well-defined task — and the open stack that made it possible is the more durable story.

OmniDocBench measures document parsing and nothing else. A score above a frontier model on this one narrow task implies nothing about general capability. Every figure below is published by the model’s own authors on their own model card. This publication has not run the benchmark and has reproduced none of them. No parent company is named here because the card names none: the organisation is XingChen-AGI and the web app sits on teleai.com.cn. The model is a rename of NaviDC-OCR, whose weights were released in August. Nothing here is investment advice.

What Happened

On the OmniDocBench v1.6 leaderboard — a benchmark that measures document parsing only, covering structured text, tables, and formulas extracted from scanned and digital pages — the TeleOCR model card reports an overall score of 96.87 for TeleOCR at 1.2 billion parameters. That places it above Ovis2.6-30B-A3B at 93.62, Gemini 3 Pro at 92.85, Gemini 3 Flash at 92.58, Qwen3-VL-235B at 89.78, GPT-5.2 at 86.52, and InternVL3.5-241B at 83.61.

Two limits belong here, immediately: OmniDocBench measures one narrow, well-defined task, and a score above GPT-5.2 on it says nothing whatsoever about general capability, reasoning, or any broader comparison. Every figure here is published by the model’s own authors on their own model card. This publication has not run the benchmark and has reproduced none of these numbers.

The sub-scores the card reports are: Text Edit 0.027, Formula CDM 96.36, Table TEDS 97.05, Table TEDS-S 98.52, and Read Order Edit 0.122. On a second table — the EMNLP 2026 Dr.DocBench Challenge — the card reports TeleOCR at 67.96 overall, against MinerU 2.5 Pro at 62.26, OvisOCR2 at 59.25, and PaddleOCR-VL-1.6 at 55.11. The card credits two methods for its performance: Multi-node Consensus Voting for automatic pseudo-label generation, and geometry-aware document modeling, with the stated goal of unifying digital and camera-captured documents in one framework — something the card says existing methods mostly handle separately.

This is not a new release. The model card’s own log records that weights and a technical report were released on 17 August 2026 under the name NaviDC-OCR, that the Hugging Face entry was created on 14 August 2026, and that a rename to TeleOCR took effect on 10 September 2026. The entry was last modified 29 September. What circulated this week is a renaming, not an unveiling.

On ownership: the organisation publishing the weights is XingChen-AGI, the web app badge points to teleai.com.cn, and the repository sits under an individual GitHub account. No parent company is identified on the model card.

The key insight: For a well-defined, high-volume document job, a purpose-built 1.2B model released under Apache-2.0 is competitive on the authors’ own benchmark with general-purpose systems that cost orders of magnitude more to run. That is not a claim that small models have caught up with frontier models. It is a claim that for tasks with a tight specification and a good benchmark, specialisation is doing work that scale is not.

Only models the card gives a parameter count for are plotted; it prints a dash for Gemini 3 Pro and GPT-5.2, w
Only models the card gives a parameter count for are plotted; it prints a dash for Gemini 3 Pro and GPT-5.2, which score 92.85 and 86.52. The benchmark measures document parsing and nothing else.

The Structural Read

The benchmark result is the headline, but the acknowledgements section is the more structurally interesting paragraph. The card states that TeleOCR is built upon MinerU, Qwen2.5-VL, Qwen3, Transformers, PyTorch, and FlashAttention. MinerU 2.5 Pro then appears in TeleOCR’s own comparison table at 95.75 — below TeleOCR’s 96.87. Building on an open project and then benchmarking against it is completely ordinary open-source practice. The licence is Apache-2.0, and the card credits those projects by name. Nothing about that is improper.

What it illustrates is how quickly an open stack compounds. A vision-language base from one lab, a parsing toolkit from another, and a specialised model on top that — on the authors’ own benchmark — reports beating both the toolkit it uses and the general models many times its size. All of this happened inside a few months, all under permissive licences. None of it has been independently verified here, and the card publishes no inference cost, no latency figure, and no adoption data beyond download and like counts.

The Product Overhang Doctrine frames this precisely. Open-weight infrastructure — base vision-language models, parsing toolkits, flash-attention kernels — accumulated permissively over the past two years. The cost to assemble a task-specific model on top of that stack dropped to the point where a team publishing under a single GitHub account could release weights that, on their own benchmark for their own task, post numbers above the largest general models.

That overhang is not exhausted. Every well-specified enterprise task sitting on top of a $10B inference bill is a candidate for the same compression.

Three Implications

The Benchmark Is the Moat, Not the Model Size

A well-specified benchmark is increasingly the competitive unit. Teams that define the task narrowly, build a labelled evaluation, and optimise against it can claim a credible leadership position at a fraction of the parameter count. For buyers of document-processing infrastructure, this means the procurement question shifts from “which frontier model?” to “which benchmark best represents my actual workload?” — a very different due-diligence exercise.

Apache-2.0 at 1.2B Is a Deployable Cost Structure

If the authors’ numbers hold under independent evaluation, a 1.2B model on Apache-2.0 is not just a research artifact — it is a deployable cost structure. Enterprises running high-volume document pipelines at current frontier-model token prices face a straightforward arithmetic case for task-specific fine-tuned alternatives. The card provides no inference cost or latency data, so that arithmetic cannot be completed here. But the parameter count alone implies an inference cost profile that general-purpose frontier APIs cannot easily match for volume workloads.

Open Stack Compounding Accelerates Provenance Complexity

TeleOCR is built on MinerU, Qwen2.5-VL, Qwen3, Transformers, PyTorch, and FlashAttention — six distinct open-source lineages assembled inside a few months by an organisation whose ownership structure is not disclosed on the model card. That is legal, and it is how open-source is supposed to work. It also means that enterprise compliance, security, and vendor-risk teams now need to evaluate not just the model they are procuring but the full dependency graph behind it.

The supply-chain question in AI is becoming structurally similar to the one the software industry has been navigating since the Log4j moment.

Business Engineer Framework

Product Overhang Doctrine — Applied to the AI Stack

The TeleOCR case is a live instance of the Product Overhang Doctrine: open-weight base models, parsing toolkits, and attention kernels accumulated permissively, and the cost to assemble a task-specific layer on top fell faster than anyone adjusted their infrastructure pricing. The Business Engineer Map of AI traces exactly where that compression is happening across the nine layers of the stack — and which positions are still exposed to the same pressure.

Read the Map of AI →
Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA