Our Hugging Face tracker has had XingChen-AGI/TeleOCR on the likes table two days running — up 16.2% yesterday to 975 likes, after 28.7% the day before — and the repo was modified yesterday. Worth a look.
It is a ~1.2B parameter vision-language model built for one job: document parsing, across both digital and camera-captured documents, in a single framework.
Why “both” is the whole point
Document OCR has quietly split into two problems that get solved separately.
Digital documents — PDFs, screenshots, exported pages — are geometrically perfect. The text is axis-aligned, evenly lit, uniformly sharp. A model trained on these is really solving a layout and character-recognition problem.
Camera-captured documents are a different beast entirely. The page is curved, because paper is not flat. It is perspective-distorted, because your phone is not perpendicular to it. It has uneven illumination and a specular highlight where the window is. It has motion blur on one side and shadow from your own hand.
Models trained on scans degrade badly on photographs, because none of that distortion is in their training distribution. The usual industrial answer is a preprocessing pipeline — detect the page, dewarp it, correct the perspective, normalise the lighting, then OCR the rectified image. That works, and it is brittle: every stage can fail, and a dewarping error produces confidently wrong text downstream.
TeleOCR’s claim is geometry-aware document modelling inside the model, so the distortion is handled as part of understanding the page rather than removed beforehand.
What’s in the technical report
Two contributions worth naming, from the report by Cai, Zou, Liu, Wang, Tang, Yang, Tong, He and Sun:
Multi-node Consensus Voting (MCV) for automatic pseudo-label generation. Training data for document parsing is expensive because annotating a page’s full structure — reading order, tables, formulas, columns — is slow, skilled work. MCV generates labels automatically by having multiple models parse the same page and keeping what they agree on. Agreement is a reasonable proxy for correctness, and disagreement flags the hard cases. It is the same logic as ensemble self-training, applied to a labelling problem where human annotation is the bottleneck.
Geometry-aware document modelling for camera captures, per above.
Plus specialised handling for tables, formulas and complex layouts — the three things that break naive OCR, because all three encode meaning in spatial relationships rather than in character sequence.
The numbers
| Benchmark | Score |
|---|---|
| OmniDocBench v1.6 | 96.87 overall |
| Wild_OmniDocBench | 88.53 overall |
| Dr.DocBench Challenge | 67.96 overall |
The gap between the first two is the interesting part: 96.87 on clean documents, 88.53 on “wild” ones. The model is genuinely better on scans than on photographs — as everything is — but an 8-point drop rather than a collapse is the claim being made, and it is the claim that matters if you intend to point a phone at things.
Dr.DocBench at 67.96 suggests the genuinely hard cases are still hard, which is the honest read.
Chinese and English. Apache 2.0 — which for a model you might want to run inside a product or an installation is the single most important line on the card.
Why a creative-technology site cares about OCR
Because it is an archive tool, and archives are raw material.
1.2B parameters runs locally. That is small enough for a consumer GPU, plausibly small enough for a well-specified laptop with quantisation. So this is not an API you send your material to — it is a model you point at a folder.
Concretely useful for:
- Digitising your own archive. Sketchbooks, notebooks, letters, printed ephemera, zines — photographed rather than scanned, because nobody is going to flatbed-scan three hundred pages. This is exactly the camera-captured case.
- Working with printed source material. Anyone making work from found text, historical documents, or public records is otherwise doing a great deal of transcription by hand.
- Feeding a generative pipeline. Text extracted from a specific corpus — a family archive, a local newspaper run, a body of technical manuals — is the input to fine-tuning, retrieval, or generative text work with an actual subject rather than a generic one.
- Accessibility, which is the unglamorous and most defensible use: making printed material in a collection readable by a screen reader.
The caveats to hold: it is Chinese and English only, benchmarks are not your documents, and handwriting is a different problem from printed text — nothing on the card claims handwriting. Test it on a representative twenty pages before you commit to a workflow around it.
One note on the version history worth knowing if you are searching for it: the model was renamed from NaviDC-OCR to TeleOCR in September 2026, and subsequent iterations continue under the TeleOCR name.