For the primary time, an OCR mannequin reads Indian languages as fluently as English paperwork. Sarvam Imaginative and prescient 2.1 breaks a long-standing trade-off. Indian companies had to decide on between structural doc parsing and script recognition. By no means each. Finance groups automating invoices confronted this. Insurance coverage firms processing handwritten claims in a number of states confronted this. Organizations digitizing regional data confronted this. All needed to choose accuracy or protection. This mannequin does each. Let’s discover what modified, the place it wins, the place it struggles, and 5 take a look at paperwork you may run your self.
Key Options of Sarvam Imaginative and prescient 2.1
Pulling named fields out of a desk moderately than transcribing the whole grid. This works for statements, ledgers, or something the place a price solely is smart in relation to its row and column headers.
The identical concept utilized to kind fields. The label and worth sit in separate bins. The pairing must be inferred from format, not studying order.
Indic Handwritten Extraction
Handwriting in Indian scripts extracted into structured fields moderately than transcribed as a block. That is the toughest functionality. Sarvam skilled it partly on video sources to seize actual handwriting selection, not simply artificial samples.
What the Numbers Truly Present
Till now, if a mannequin was good at English doc parsing it was often mediocre at Indian languages, and the fashions that dealt with Indic scripts nicely weren’t aggressive on common doc construction. These have been separate instruments with separate failure modes.

Every level is one system scored on each axes. Higher proper is best on each directly.
Have a look at the place the opposite factors sit. Infinity-Parser2 Professional scores 86.1 on English, second solely to Sarvam, after which collapses to 49.83 on Indic. Google Cloud Imaginative and prescient does the reverse: 39.6 on English doc construction, however 71.76 on Indian languages, as a result of it has had Indic OCR for years with out the format intelligence. Sarvam 2.1 is the one level within the higher proper.
The place it Wins and The place it Doesn’t
The Indic Benchmark
Sarvam additionally launched the benchmark itself: 6,909 samples, 6,609 spanning all 22 official Indian languages and 300 in English, drawn from newspapers, brochures, textbooks and historic writing dated from 1800 to the current.
The benchmark is published on Hugging Face, which issues. A vendor-built benchmark that the seller wins is value little by itself. A vendor-built benchmark launched publicly so others can run it’s a totally different proposition, and it’s the proper approach to do that.
olmOCR-Bench
The group benchmark for English doc parsing. It runs pass-fail checks throughout eight classes: arXiv maths, base textual content, headers and footers, tiny textual content, multi-column pages, degraded outdated scans, old-scan maths, and tables. It assessments whether or not information arepresent or absent moderately than scoring delicate variations.
Sarvam notes that this benchmark is formally English-only however incorporates some contaminant samples in Chinese language and different scripts. Their earlier launch reported on a filtered English-only set; this time they report on the official set for parity with opponents, which is the extra conservative selection.
| Mannequin | Math | Tables | OldScan | MultCol | General |
| Sarvam Imaginative and prescient 2.1 | 90.5 | 91.9 | 55.3 | 82.1 | 87.3 |
| Infinity-Parser2 Professional | 87.4 | 88.9 | 58.0 | 83.3 | 86.1 |
| Opus 5 | 90.0 | 89.5 | 54.0 | 85.8 | 85.1 |
| Chandra-OCR2 | 86.5 | 87.5 | 49.2 | 82.4 | 84.5 |
| Mistral OCR4 | 83.7 | 88.6 | 48.9 | 85.7 | 83.1 |
| Gemini 3.6 Flash | 86.5 | 85.9 | 48.1 | 78.6 | 82.4 |
| GPT 6 Astra | 82.6 | 90.9 | 47.0 | 77.8 | 81.8 |
OmniDocBench v1.6
A distinct measure: structural constancy moderately than truth presence. It’s a composite of textual content edit distance, desk construction scored with TEDS, formulation recognition scored with CDM, and studying order, run over newspapers, textbooks, magazines and monetary experiences.
| Mannequin | Textual content edit dist (decrease higher) | Method CDM | Desk TEDS | General |
| PaddleOCR-VL 1.6 | 0.0356 | 0.985 | 0.931 | 96.01 |
| Sarvam Imaginative and prescient 2.1 | 0.0289 | 0.988 | 0.890 | 94.97 |
| GLM-OCR | 0.0374 | 0.984 | 0.895 | 94.71 |
| GPT 6 Astra | 0.0460 | 0.967 | 0.891 | 93.74 |
| Gemini 3.6 Flash | 0.0371 | 0.976 | 0.869 | 93.58 |
| Opus 5 | 0.0471 | 0.967 | 0.856 | 92.51 |
Learn that desk throughout moderately than down. Sarvam has the very best textual content edit distance of any mannequin at 0.0289 and the very best formulation rating at 0.988. It loses the highest spot purely on desk construction, the place PaddleOCR-VL scores 0.931 towards Sarvam’s 0.890. Sarvam calls each benchmarks arguably saturated, which is truthful when the highest twelve fashions sit inside 4 factors of one another.
On the International Benchmarks
Sarvam leads olmOCR-Bench at 87.3 general. It doesn’t lead each class:
| Class | Sarvam 2.1 | Finest rating | Held by |
| Outdated scans | 55.3 | 58.0 | Infinity-Parser2 Professional |
| Multi-column | 82.1 | 85.8 | Opus 5 |
| Tiny textual content | 92.5 | 93.5 | Opus 5 |
| Tables | 91.9 | 91.9 | Sarvam 2.1 |
| Math | 90.5 | 90.5 | Sarvam 2.1 |
And on OmniDocBench v1.6 Sarvam is second, not first: 94.97 towards PaddleOCR-VL 1.6 at 96.01. Sarvam wins on textual content edit distance and formulation recognition, PaddleOCR wins on desk construction with a TEDS of 0.931 towards 0.890.
Santhali is the clear loss. Sarvam scores 53.91 and Bodhan Indic-OCR scores 68.30, a spot of greater than fourteen factors. Odia is a narrower loss to Gemini 3.6 Flash, 80.01 towards 81.01. Kashmiri just isn’t a loss however it’s weak in absolute phrases at 54.82, the very best rating any mannequin manages on that language.
Sarvam Imaginative and prescient 2.1 Structure

- The vision-language mannequin doesn’t learn the web page instantly
- Two harnesses sit in entrance of it: a semantic format parser that segments the web page into areas, and a pointer community that establishes studying order
- The VLM can work at web page stage alone, however Sarvam discovered the accuracy trade-off makes harnessing worthwhile
- That is why multi-column newspapers and merged-cell tables are the arduous take a look at instances: if the harness segments wrongly, the VLM transcribes appropriate textual content within the improper order. Fluent and improper, which is tougher to catch than garbled output
- Publish-training mixed supervised fine-tuning with RLVR (reinforcement studying with verifiable rewards). This suits OCR nicely since correctness towards a identified transcription is programmatically checkable.
Learn how to Run Sarvam Imaginative and prescient 2.1
I couldn’t run these myself. Sarvam’s API isn’t reachable from the atmosphere I work in, so each end result has to return from you. What I’ve achieved as a substitute is construct the paperwork, write the precise floor fact for every, and mark the particular failure to look at for. That turns a imprecise take a look at this right into a scoreable take a look at that takes about fifteen minutes.
The quickest route is the document intelligence playground, which wants no code. Add, run, evaluate towards the reply key.
For the API, there are two endpoints and so they do totally different jobs:
# Digitise: full-page conversion to structured textual content with format preserved
# use for paperwork 1, 4 and 5
POST
# Extract: key-value pairs, tables, kind fields
# use for paperwork 2 and three
POST
Run doc 2 by each. The distinction between what digitise returns and what extract returns on the identical kind is the clearest demonstration of what the brand new extraction functionality truly provides.
Conclusion
Sarvam 2.1 is the primary mannequin that doesn’t drive a selection between English doc construction and Indian language protection. That’s an actual end result. It’s the appropriate match for Indian-language paperwork, printed or handwritten, types and tables needing structured extraction, mixed-script pages, and manufacturing pipelines that want predictable price.
It’s not the appropriate match if Santhali or Kashmiri is your major language, if desk construction constancy is non-negotiable, or in case your paperwork are closely degraded historic scans, the place no mannequin performs nicely but.
The mannequin can be nonetheless 55.3 on outdated scans and 53.91 on Santhali. These two numbers will determine whether or not it really works to your paperwork, not the headline.
Login to proceed studying and revel in expert-curated content material.
