How a compile is measured.

The reader study compares 10 models on the same 1,651 benchmark pages. The recovery study below tests our own pipeline with and without one lane. Both publish methods and limits; neither is a current-service score.

TAVONEL research · same-pipeline comparison

Recovery makes a measurable difference on difficult pages.

80.6with recovery
53.7without recovery

olmOCR-Bench score across 1,403 documents and 8,413 checks, measured 2026-08-08. This isolates one recovery lane in a historical research pipeline. It is not a current-service score or a competitor comparison. Low-quality old scans reached 36.9% even with recovery.

Method and limits

On olmOCR-Bench, scored by the benchmark's own evaluator at revision cfa88c1e, the same pipeline scored 80.6 with the recovery lane and 53.7 with only that lane switched off — a gap of 26.9 points, 95% confidence intervals 79.62–81.57 and 52.62–54.93, which do not overlap. Model, evaluator revision, corpus, source manifest, test set and settings were identical; the only difference was whether the documents recovery delivered carried their content. Measured 2026-08-08 over 1,403 documents and 8,413 checks. One category, headers and footers, scores higher without recovery, because a check that a phrase is absent passes trivially on an empty page — the no-recovery figure is generous rather than harsh.

Download the R-01 result and receipt

Document reading, measured here — Model Arena, 2026-09-06 19:55 KST

These scores measure page reading only. Every row is a model we ran ourselves on OmniDocBench, scored by OmniDocBench end2end quick_match at evaluator revision 193627ae9e97, under the conditions listed below.

Text Edit distance against the benchmark’s ground truth

Lower is better · normalised edit distance, 0 to 1 · bars are drawn to the largest figure on this chart, 0.1790, not to 1.0.

  1. OvisOCR20.02901,651 pages
  2. PaddleOCR-VL 1.60.04261,651 pages
  3. HPD-Parsing0.04291,651 pages
  4. Infinity-Parser2-Flash0.04491,651 pages
  5. MinerU VLM (MinerU2.5-Pro-2605-1.2B)0.05081,651 pages
  6. MonkeyOCRv2-B-Parsing0.05781,651 pages
  7. DeepSeek-OCR-20.05841,651 pages
  8. MinerU 3.4.5 Pipeline (PP-OCRv6)0.05901,651 pages
  9. Unlimited-OCR0.09791,651 pages
  10. olmOCR-2-7B-1025-FP80.17901,651 pages

Table structure similarity (TEDS)

Higher is better · tree-edit-distance similarity, 0 to 1 · bars are drawn to the largest figure on this chart, 0.9344, not to 1.0.

  1. OvisOCR20.93421,651 pages
  2. PaddleOCR-VL 1.60.93441,651 pages
  3. HPD-Parsing0.85141,651 pages
  4. Infinity-Parser2-Flash0.82021,651 pages
  5. MinerU VLM (MinerU2.5-Pro-2605-1.2B)0.92651,651 pages
  6. MonkeyOCRv2-B-Parsing0.82431,651 pages
  7. DeepSeek-OCR-20.78461,651 pages
  8. MinerU 3.4.5 Pipeline (PP-OCRv6)0.79641,651 pages
  9. Unlimited-OCR0.86231,651 pages
  10. olmOCR-2-7B-1025-FP80.56031,651 pages

Median seconds per page against text Edit distance

Lower is better on both axes, so the bottom-left corner is the cheapest reading per page at the closest match to the benchmark’s ground truth. Every dot is one model on the same corpus, the same evaluator revision and the same scoring driver.

OvisOCR2: 4 s/page (median over 5,132 timed pages), text Edit 0.0290 over 1,651 pagesPaddleOCR-VL 1.6: 4 s/page (median over 5,117 timed pages), text Edit 0.0426 over 1,651 pagesHPD-Parsing: 3 s/page (median over 5,117 timed pages), text Edit 0.0429 over 1,651 pagesMinerU VLM (MinerU2.5-Pro-2605-1.2B): 5 s/page (median over 5,117 timed pages), text Edit 0.0508 over 1,651 pagesMonkeyOCRv2-B-Parsing: 6 s/page (median over 5,093 timed pages), text Edit 0.0578 over 1,651 pagesDeepSeek-OCR-2: 27 s/page (median over 5,132 timed pages), text Edit 0.0584 over 1,651 pagesMinerU 3.4.5 Pipeline (PP-OCRv6): 4 s/page (median over 5,132 timed pages), text Edit 0.0590 over 1,651 pagesUnlimited-OCR: 29 s/page (median over 5,107 timed pages), text Edit 0.0979 over 1,651 pagesolmOCR-2-7B-1025-FP8: 13 s/page (median over 5,132 timed pages), text Edit 0.1790 over 1,651 pages029 s/page0.17900
  1. OvisOCR2 4 s/page (900 pages/h serial, median over 5,132 timed pages) · text Edit 0.0290 over 1,651 pages
  2. PaddleOCR-VL 1.6 4 s/page (900 pages/h serial, median over 5,117 timed pages) · text Edit 0.0426 over 1,651 pages
  3. HPD-Parsing 3 s/page (1,200 pages/h serial, median over 5,117 timed pages) · text Edit 0.0429 over 1,651 pages
  4. MinerU VLM (MinerU2.5-Pro-2605-1.2B) 5 s/page (720 pages/h serial, median over 5,117 timed pages) · text Edit 0.0508 over 1,651 pages
  5. MonkeyOCRv2-B-Parsing 6 s/page (600 pages/h serial, median over 5,093 timed pages) · text Edit 0.0578 over 1,651 pages
  6. DeepSeek-OCR-2 27 s/page (133.3 pages/h serial, median over 5,132 timed pages) · text Edit 0.0584 over 1,651 pages
  7. MinerU 3.4.5 Pipeline (PP-OCRv6) 4 s/page (900 pages/h serial, median over 5,132 timed pages) · text Edit 0.0590 over 1,651 pages
  8. Unlimited-OCR 29 s/page (124.1 pages/h serial, median over 5,107 timed pages) · text Edit 0.0979 over 1,651 pages
  9. olmOCR-2-7B-1025-FP8 13 s/page (276.9 pages/h serial, median over 5,132 timed pages) · text Edit 0.1790 over 1,651 pages

Rows we did not rank

The hosted row. opus5_subscription ran on a subscription surface, not on a GPU we rented, so it has no hardware line. Its list-price reference on 2026-09-03 was $5/Mtok in and $25/Mtok out. masterplan section 4 records this as an API list-price reference. This run uses the subscription surface, so the number is only used for api_equivalent_list_price_usd. Never report the subscription lane as $0/page (ARENA_CONTRACT section 7).

Not supported

Blind quality detection failed

We tested whether prediction-only signals could pick the worst documents without ground truth. They could not beat ranking by length alone. Published as unsupported, and not shipped as a feature.

Nothing on this page, and no routing decision behind it, reads a scalar quality score. Read the finding.

Conditions, licence and pins

  • Dataset. OmniDocBench at revision aa1ee96d106d — licensed research-only-non-commercial, redistribution prohibited without separate rights review. Its pages are not republished here and no page image from it appears on this site.
  • Evaluator. OmniDocBench end2end quick_match from opendatalab/OmniDocBench at pin 193627ae9e97d89188468ed1ee3b7a856ff76044, licence Apache-2.0, entrypoint python pdf_validation.py --config config.yaml. Scoring driver: reports/full_compare_20260905/_win_omnidoc_score_full.py.
  • Speed. Inference latency per page from the campaign database (SUCCESS started_at to finished_at), as a median and the serial pages per hour that median implies. Scoring wall-time. OmniDoc elapsed and olmOCR scoring time are not speed.
  • Registry snapshot. Model revisions and the GPU catalogue were resolved at 2026-09-05T03:01:01Z. A listed GPU rate is the provider’s rate when that pod was provisioned; it is raw hardware cost and never a price for a page.

Limits of this run

  • Speed = inference latency from campaign.sqlite (median sec/page, serial pages/h). NOT scoring wall-time.
  • Do not use OmniDoc elapsed_min or olmOCR score_min as speed.
  • ParseBench layout = N/A for all models (elements_available=false).
  • opus5_subscription and infinity_parser2_flash missing from speed board.
  • infinity_parser2_pro is FOUNDER_EXCLUDED — reference-only (speed shown; quality pending/excluded).
  • glm_ocr OmniDoc remains REF until optional SUCCESS-only rescore.
  • monkeyocrv2_b OmniDoc restored home-native text Edit ~0.0578.
  • See ANOMALY_NOTES.md.
Receipts: model identity, hardware, and the digest of every source
Identity and hardware per row. A blank cell is a value the campaign did not record; it is never filled from another row.
ModelWeights or hosted idRevisionLicenceGPUListed ratePrice snapshot
OvisOCR2ovisocr2ATH-MaaS/OvisOCR21fc9221b7823a371d6e97f92d527cc847e24e107apache-2.0NVIDIA GeForce RTX 4090$0.74/h listedsha256:7cc1f10400f2b4d675c9a3dd91cafc9532f1f77c4a885ab2da911c0651e05cd0
PaddleOCR-VL 1.6paddleocr_vl_1_6PaddlePaddle/PaddleOCR-VL-1.6c5630abae1d940eafe0697512a0325494b02ab42apache-2.0NVIDIA GeForce RTX 4090$0.74/h listedsha256:a6d8a7deb0f96cb419712746b451078c4961fdeace70e65dc928f9318d58777f
HPD-Parsinghpd_parsingPaddlePaddle/HPD-Parsing91de80054c23ab4238e3d7073fa2b83c2a7e301eapache-2.0NVIDIA H100 80GB HBM3$3.49/h listedsha256:1a61d52222faa9c0ddad4a257ab6e70491adaf93ebb8b4abedd6a0483494aba7
Infinity-Parser2-Flashinfinity_parser2_flashinfly/Infinity-Parser2-Flashcfddc4106b0abc4706f05575a164f7dd35d09c46apache-2.0——no canary proof in the campaign evidence directory for this model key
MinerU VLM (MinerU2.5-Pro-2605-1.2B)mineru_vlmopendatalab/MinerU2.5-Pro-2605-1.2Bbff20d4ae2bf202df9f45284b4d43681555a97edLicenseRef-MinerU-Open-Source-LicenseNVIDIA A40$0.49/h listedsha256:d740783489bd956adb72c9fa72d2613b47e406121136a96f5c3022013e272987
MonkeyOCRv2-B-Parsingmonkeyocrv2_bzenosai/MonkeyOCRv2-B-Parsing2419139b7bcd3fda2689b2a83167172afba91c8bapache-2.0NVIDIA GeForce RTX 4090$0.74/h listedsha256:139817ec531fe88c84c8c2ccf05e2dcf96b3b21a28cceb0f391b7676ef1df95a
DeepSeek-OCR-2deepseek_ocr2deepseek-ai/DeepSeek-OCR-2aaa02f3811945a91062062994c5c4a3f4c0af2b0apache-2.0NVIDIA GeForce RTX 4090$0.74/h listedsha256:fd217ef9947756c2a70a6a9dc5fc494ecb2bd5cabc4958de098b74464626734a
MinerU 3.4.5 Pipeline (PP-OCRv6)mineru_pipelineopendatalab/PDF-Extract-Kit-1.0ed6b654c018d742e65a17671e379c5e6ecc87ec9LicenseRef-MinerU-Open-Source-LicenseNVIDIA GeForce RTX 4090$0.74/h listedsha256:dc9357bc0c9b24ceda6198706ac7f83c7d56dad5ac59304cdb23f91ceb5d6b12
Unlimited-OCRunlimited_ocrbaidu/Unlimited-OCR07dea832e22aefee32ad281d4b80551282e1c168mitNVIDIA GeForce RTX 4090$0.74/h listedsha256:5e8ed330efe3d8e378ed4c2f07e36f97ca235835d7e18394b2851d48c72542e6
olmOCR-2-7B-1025-FP8olmocr2allenai/olmOCR-2-7B-1025-FP840bd7202494b8264ee17ada08b401b5aab7a9ce1apache-2.0NVIDIA GeForce RTX 4090$0.74/h listedsha256:342c3e2a45092d6bc3c0efa5aac98deb1738c2ee568bfb84acf2dd69107dfdf8
GLM-OCRglm_ocrzai-org/GLM-OCRca5d8b3e287e52589e37c28385d9655ee4372f9dmitNVIDIA GeForce RTX 4090$0.74/h listedsha256:1230a3f269dc3772d0c8c6403afe8a2f537cb6a66832ee78958cb74146c8720e
Claude Opus 5 - Claude Code subscription surfaceopus5_subscriptionclaude-opus-5—proprietary-anthropic-commercial-terms——no canary proof in the campaign evidence directory for this model key
Infinity-Parser2-Proinfinity_parser2_proinfly/Infinity-Parser2-Prob27d470100514329fc6439aada8f16ccea5f9e2aapache-2.0NVIDIA H100 80GB HBM3$3.49/h listedsha256:e9f8cf950b1e243ec7b1156c1754d867f8b7e3f9bcc88eb3131ab04adfaf73de
The campaign files this board was assembled from, and the sha256 of the bytes the build read. A source that is not on the build machine stops the build.
Sourcesha256
evaluator_registry.json668eaa0fb6fc7f823ac88cb20b692dc4d45a2c114fd4d7a6e0f5457b454a380b
evidence/canary-proof-deepseek_ocr2.json011ea7c988a31dfe66b7caa7ba88f34f246179d201028767c50d3ea1d51931c9
evidence/canary-proof-glm_ocr.jsonfc30bde2b38d20c363179acdfcfc825c580f7411b2c4ab1492fb54354fa1fef8
evidence/canary-proof-hpd_parsing.json9ed6340270ec6ff7c9aa037f7b5258b02b1a53d91daef0f6d11d2d079e2e334b
evidence/canary-proof-infinity_parser2_pro.jsonb5f5a1e1fc62193312b9c86e6988e5b78627a3b9cf9f5831d47a403d9ae378a9
evidence/canary-proof-mineru_pipeline.json8f82ecc4a49cfac6e12d966aa69bb845fac887b495da85dd3aac2acd2b7f9835
evidence/canary-proof-mineru_vlm.json07eba250b9b61cb8f9b024fc6ed8c0a411f4f3015b3e6a68ec309365e7945426
evidence/canary-proof-monkeyocrv2_b.jsoncebd226979599ce36b94713986c3acf7ea20fcd5c5a4038bdd2d468e0457ab89
evidence/canary-proof-olmocr2.json6b10ab06116cf7d6eb0eddc6a37980c388328bde2bcce695deaf58802c4d8896
evidence/canary-proof-ovisocr2.json595ee2db9cb47289ce7473ec84d4bff792b227c619a6358a6458f4f923ca3389
evidence/canary-proof-paddleocr_vl_1_6.jsonc69e60c938e85091851ad53dd9af13156668f5d53f82c76839882b64dff1e87a
evidence/canary-proof-unlimited_ocr.json6409f0204be62c9db95b1030e2d99dcf36a046cd29dd60013e4671516bfe13dc
model_registry.json83152ef92e2792a7ab3b2ef25fa9774115adfea2ebcd45363f7207764f8ee693
reports/full_compare_20260905/ANOMALY_NOTES.mdb7c64194f638a9a2ccbe3c38a2484d2f6d949a504e9a1cc1926cdfb23b44da9d
reports/full_compare_20260905/LEADERBOARD_QUALITY_SPEED.json2a699cbf8c107bb57a4f837a9e053028a2cbaf18007a612f26ef0596d504f260
reports/full_compare_20260905/STATUS.mdbf2e367a6f2bf4189f00e090713164e0e68daf7c20dbeeb1c72ef2b107e652ea
reports/full_compare_20260905/comparison_omnidoc_full.jsond95eefd4bae67cf15ee975dac327a4b5bd533f07d01d43c6d1a9967fd82a49e5

The other run on this site

The qualification contract defines what a knowledge-compilation result has to carry: a frozen configuration, corpus and output digests, a named denominator, published failures, and reproducible scoring material.

GDP.pdf evaluation design

GDP.pdf is a public set of 100 held-out tasks across ten professional domains. Our planned evaluation keeps the corpus, prompts, target models, and scoring rubric fixed across four arms. Results will be published here only with a qualified run receipt.

Review the public dataset and reference harness.

Read the two-arm run: compiled context against the PDF alone — a first result on this design, with its deviations, denominators and receipts. It is not comparable with any published GDP.pdf score.

Published supporting research

These findings remain scoped to the population and question named in each downloadable receipt.

The eight metric families

Document reading

What the reader recovered from the page, and what it invented or dropped: text, layout, tables, formulas, reading order, and the coordinates every later stage binds to.

  • text fidelity to the source
  • layout
  • table
  • formula
  • reading order
  • region placement
  • hallucinated and omitted content
  • latency
  • GPU memory
  • cost

Evidence

Whether each compiled statement is bound to the source region that actually supports it, and how much of the world is bound at all.

  • evidence binding precision
  • evidence coverage
  • source citation exactness
  • region correctness

Identity

Whether two mentions of one thing became one object, and at what cost in wrong merges, wrong splits, and cases left open for a person.

  • entity resolution precision
  • false merge
  • false split
  • unresolved rate

Knowledge

Whether the objects, claims and relations built on top of the read are correct, valid against the schema, and honest about the places sources disagree.

  • claim correctness
  • relation correctness
  • ontology validity
  • conflict detection

Temporal

Whether the world knows which revision it is holding: what has been superseded, what has gone stale, and what a past answer stood on at the time.

  • stale knowledge rate
  • supersession correctness
  • point-in-time correctness

Recompilation

Whether a source change was traced to exactly the parts it invalidated — no wider, and crucially no narrower — and what that saved against rebuilding everything.

  • affected set precision
  • affected set recall
  • work avoided
  • cost avoided
  • equivalence with a full rebuild
  • publish refusal correctness

Ask

Whether an answer stands on cited source, points at the right region, and declines when the world cannot support it.

  • grounded answer precision
  • citation precision
  • abstention correctness
  • stale answer prevention

Operations

What it costs and how long it takes for a change in a source to reach a world an agent is allowed to read.

  • median and 95th-percentile time from source change to Active World
  • cost per 1,000 pages
  • cost per changed semantic unit

What a result has to carry

Every figure published here binds to a record, and five things have to be true of it before the figure may be read as a result.

The complete receipt schema

A record missing a digest, or missing the population a rate was measured over, is refused by the build rather than rendered with a blank cell. These are the fields the validator checks.

Dataset
Which corpus was read.
Dataset version
Which cut of it, because corpora are edited.
Corpus digestsha256
sha256 over the corpus, so “the same dataset” is a checkable statement rather than a claim.
Denominator
The population every rate in the record was measured over, and how many members it has.
Model id
The exact model, never the vendor or the family name.
Model revision
The revision served on the day. Two runs of “the same model” are not the same run.
Input mode
What the model was actually given: the page image, extracted text, or both.
Prompt digestsha256
sha256 of the prompt, frozen before the run started.
Config digestsha256
sha256 of the run configuration: schema, decoding, retries, thresholds.
Compiler version
The TAVONEL build that compiled the world the metrics were read from.
Hardware
The machine. Latency and cost are properties of a machine, not of a model.
GPU
The accelerator, or “none” for a CPU run. The field is answered, never left blank.
Runtime
The serving runtime and its version.
Price snapshot
The prices in force on the day, because every cost figure expires.
Raw result digestsha256
sha256 of the unaggregated predictions, so an average can be recomputed by someone else.
World digestsha256
sha256 of the compiled world the metrics were read from.
Run receipt digestsha256
sha256 of the receipt that ties all of the above to one execution.
Date
The day the run happened, as YYYY-MM-DD.
Metrics
Every figure, with its family, its unit, and its own denominator.
Comparison basis
same_condition for a run we executed; quoted for someone else’s published score, which stays theirs.
Published failures
What the run got wrong, published with it.

Qualification rules

North Star metric

Definition

Verified Fresh Knowledge Coverage

The share of knowledge in an Active World that carries source evidence, passes validation, and agrees with the latest revision of the source it came from.

Supporting metrics

  • Evidence Coverage
  • Stale Knowledge Rate
  • Recompile Avoidance
  • Equivalence Pass Rate
  • Identity Review Rate
  • Conflict Rate
  • 95th-percentile time from source change to Active World
  • Grounded Ask Precision
  • Correct Abstention Rate

Also in the trust case: Research notes · Evidence · Reproducibility · Research · Trust

Next: Reproducibility

Can I rebuild it myself?