The reader study compares 10 models on the same 1,651 benchmark pages. The recovery study below tests our own pipeline with and without one lane. Both publish methods and limits; neither is a current-service score.
TAVONEL research · same-pipeline comparison
Recovery makes a measurable difference on difficult pages.
80.6with recovery
53.7without recovery
olmOCR-Bench score across 1,403 documents and 8,413 checks, measured 2026-08-08. This isolates one recovery lane in a historical research pipeline. It is not a current-service score or a competitor comparison. Low-quality old scans reached 36.9% even with recovery.
Method and limits
On olmOCR-Bench, scored by the benchmark's own evaluator at revision cfa88c1e, the same pipeline scored 80.6 with the recovery lane and 53.7 with only that lane switched off — a gap of 26.9 points, 95% confidence intervals 79.62–81.57 and 52.62–54.93, which do not overlap. Model, evaluator revision, corpus, source manifest, test set and settings were identical; the only difference was whether the documents recovery delivered carried their content. Measured 2026-08-08 over 1,403 documents and 8,413 checks. One category, headers and footers, scores higher without recovery, because a check that a phrase is absent passes trivially on an empty page — the no-recovery figure is generous rather than harsh.
Document reading, measured here — Model Arena, 2026-09-06 19:55 KST
These scores measure page reading only. Every row is a model we ran ourselves on OmniDocBench, scored by OmniDocBench end2end quick_match at evaluator revision 193627ae9e97, under the conditions listed below.
Text Edit distance against the benchmark’s ground truth
Lower is better · normalised edit distance, 0 to 1 · bars are drawn to the largest figure on this chart, 0.1790, not to 1.0.
Median seconds per page against text Edit distance
Lower is better on both axes, so the bottom-left corner is the cheapest reading per page at the closest match to the benchmark’s ground truth. Every dot is one model on the same corpus, the same evaluator revision and the same scoring driver.
OvisOCR24 s/page (900 pages/h serial, median over 5,132 timed pages) · text Edit 0.0290 over 1,651 pages
PaddleOCR-VL 1.64 s/page (900 pages/h serial, median over 5,117 timed pages) · text Edit 0.0426 over 1,651 pages
HPD-Parsing3 s/page (1,200 pages/h serial, median over 5,117 timed pages) · text Edit 0.0429 over 1,651 pages
MinerU VLM (MinerU2.5-Pro-2605-1.2B)5 s/page (720 pages/h serial, median over 5,117 timed pages) · text Edit 0.0508 over 1,651 pages
MonkeyOCRv2-B-Parsing6 s/page (600 pages/h serial, median over 5,093 timed pages) · text Edit 0.0578 over 1,651 pages
DeepSeek-OCR-227 s/page (133.3 pages/h serial, median over 5,132 timed pages) · text Edit 0.0584 over 1,651 pages
MinerU 3.4.5 Pipeline (PP-OCRv6)4 s/page (900 pages/h serial, median over 5,132 timed pages) · text Edit 0.0590 over 1,651 pages
Unlimited-OCR29 s/page (124.1 pages/h serial, median over 5,107 timed pages) · text Edit 0.0979 over 1,651 pages
olmOCR-2-7B-1025-FP813 s/page (276.9 pages/h serial, median over 5,132 timed pages) · text Edit 0.1790 over 1,651 pages
Rows we did not rank
reference
GLM-OCR
Reference row: on the board for context, not ranked with the settled rows.
text Edit 0.0444 · table TEDS 0.4941 · reading-order Edit 0.1796 · over 1,599 pages · tag box_success_only_after_failed_inject
glm_ocr OmniDoc SUCCESS-only rescore (2026-09-06 20:27 KST)
- **Kept REF** (founder policy / `_build_comparison.py` REF set); board REF row updated.
- BEFORE (pre-inject, 47 empties in 1651 GT match): text Edit **0.0846**, formula **0.2960**, table Edit **0.5014**, TEDS **0.4874**, reading **0.2127**, pages **1651**, tag `settled_full_1651_parallel`.
- AFTER (box CPU workers=2, $0): text Edit **0.0444**, formula **0.2708**, table Edit **0.4914**, TEDS **0.4941**, reading **0.1796**, pages **1599**, tag `box_success_only_after_failed_inject`, elapsed **612.9s**, returncode=0.
- Prepare: 1599 nonempty SUCCESS-only; 52 FAILED excluded (47 empty+5 missing).
- GT: filtered OmniDocBench to 1599 SUCCESS-matched pages (full bench source; missing FAILED would otherwise score as empty under full 1651 GT).
- Note: initial zip extract via system unzip corrupted CJK filenames (90 orphans); fixed via Python UTF-8 extract before final score.
- Soft-skip installed: `omnidoc_raw/glm_ocr/` and `C:/Users/yspow/box_omnidoc_results/`.
- Prior REF metrics backed up under `omnidoc_raw/glm_ocr/BACKUP_pre_success_only/`.
reference
Claude Opus 5 - Claude Code subscription surface
Reference row: on the board for context, not ranked with the settled rows.
text Edit 0.0896 · table TEDS 0.8284 · reading-order Edit 0.1675 · over 1,651 pages · tag settled_full_1651_parallel
Why, in the campaign’s own words
REF; 21 prepare-skipped under full GT; speed missing from speed board
opus5_subscription ParseBench table ≈ 0
- Cause: GFM pipe tables vs scorer requiring HTML `<table>`.
- Fix path: pipe→HTML re-prepare + table-only rescore (no re-inference).
opus5_subscription olmOCR timing
- **overall 0.740372 is valid.**
- Recorded score_min ~0.2 min is anomalous (likely resume/cache) — **do not use for speed ranking.**
excluded
Infinity-Parser2-Pro
FOUNDER_EXCLUDED 2026-09-04 over GPU spend (H100x2, about $7/h per pod). Speed is reference-only; no quality reading was completed, so this model has no quality row.
The hosted row. opus5_subscription ran on a subscription surface, not on a GPU we rented, so it has no hardware line. Its list-price reference on 2026-09-03 was $5/Mtok in and $25/Mtok out. masterplan section 4 records this as an API list-price reference. This run uses the subscription surface, so the number is only used for api_equivalent_list_price_usd. Never report the subscription lane as $0/page (ARENA_CONTRACT section 7).
Not supported
Blind quality detection failed
We tested whether prediction-only signals could pick the worst documents without ground truth. They could not beat ranking by length alone. Published as unsupported, and not shipped as a feature.
Nothing on this page, and no routing decision behind it, reads a scalar quality score. Read the finding.
Conditions, licence and pins
Dataset. OmniDocBench at revision aa1ee96d106d — licensed research-only-non-commercial, redistribution prohibited without separate rights review. Its pages are not republished here and no page image from it appears on this site.
Evaluator. OmniDocBench end2end quick_match from opendatalab/OmniDocBench at pin 193627ae9e97d89188468ed1ee3b7a856ff76044, licence Apache-2.0, entrypoint python pdf_validation.py --config config.yaml. Scoring driver: reports/full_compare_20260905/_win_omnidoc_score_full.py.
Speed. Inference latency per page from the campaign database (SUCCESS started_at to finished_at), as a median and the serial pages per hour that median implies. Scoring wall-time. OmniDoc elapsed and olmOCR scoring time are not speed.
Registry snapshot. Model revisions and the GPU catalogue were resolved at 2026-09-05T03:01:01Z. A listed GPU rate is the provider’s rate when that pod was provisioned; it is raw hardware cost and never a price for a page.
Limits of this run
Speed = inference latency from campaign.sqlite (median sec/page, serial pages/h). NOT scoring wall-time.
Do not use OmniDoc elapsed_min or olmOCR score_min as speed.
ParseBench layout = N/A for all models (elements_available=false).
opus5_subscription and infinity_parser2_flash missing from speed board.
infinity_parser2_pro is FOUNDER_EXCLUDED — reference-only (speed shown; quality pending/excluded).
glm_ocr OmniDoc remains REF until optional SUCCESS-only rescore.
monkeyocrv2_b OmniDoc restored home-native text Edit ~0.0578.
See ANOMALY_NOTES.md.
Receipts: model identity, hardware, and the digest of every source
Identity and hardware per row. A blank cell is a value the campaign did not record; it is never filled from another row.
The campaign files this board was assembled from, and the sha256 of the bytes the build read. A source that is not on the build machine stops the build.
Does compiled context change what a model gets right?
A held-out task set answered twice by the same model — once from the PDF alone, once with TAVONEL’s compiled context beside it. The measured delta, the conditions it was measured under and the deviations from the sealed protocol are on its own page.
The board above. It answers what a page costs to read and how closely each reader reproduced it, which is the evidence a routing policy is learned from.
The qualification contract defines what a knowledge-compilation result has to carry: a frozen configuration, corpus and output digests, a named denominator, published failures, and reproducible scoring material.
GDP.pdf evaluation design
GDP.pdf is a public set of 100 held-out tasks across ten professional domains. Our planned evaluation keeps the corpus, prompts, target models, and scoring rubric fixed across four arms. Results will be published here only with a qualified run receipt.
The provider receives the native PDF and the benchmark prompt, with no added context.
Compiled context
The same native PDF and prompt are paired with TAVONEL's sealed compiled context. This is the primary comparison with Native PDF.
Fixed retrieval
The same compiled corpus is queried through a fixed embedder and reranker, then cited evidence is passed to the same target model.
Adaptive routing
The same corpus and query pass through a sealed eligible routing policy. The receipt records the chosen candidate and, where available, its control counterfactual or shadow run.
Published supporting research
These findings remain scoped to the population and question named in each downloadable receipt.
2026-08-08
Recovery changes the outcome
What this means for you: on a page that is hard to read, the recovery step is the difference between getting the document's content and getting an empty page.
What this means for you: when a document points at something that was not supplied with it, the compiler refuses that document rather than emitting a world with a dead link in it.
What the reader recovered from the page, and what it invented or dropped: text, layout, tables, formulas, reading order, and the coordinates every later stage binds to.
text fidelity to the source
layout
table
formula
reading order
region placement
hallucinated and omitted content
latency
GPU memory
cost
Evidence
Whether each compiled statement is bound to the source region that actually supports it, and how much of the world is bound at all.
evidence binding precision
evidence coverage
source citation exactness
region correctness
Identity
Whether two mentions of one thing became one object, and at what cost in wrong merges, wrong splits, and cases left open for a person.
entity resolution precision
false merge
false split
unresolved rate
Knowledge
Whether the objects, claims and relations built on top of the read are correct, valid against the schema, and honest about the places sources disagree.
claim correctness
relation correctness
ontology validity
conflict detection
Temporal
Whether the world knows which revision it is holding: what has been superseded, what has gone stale, and what a past answer stood on at the time.
stale knowledge rate
supersession correctness
point-in-time correctness
Recompilation
Whether a source change was traced to exactly the parts it invalidated — no wider, and crucially no narrower — and what that saved against rebuilding everything.
affected set precision
affected set recall
work avoided
cost avoided
equivalence with a full rebuild
publish refusal correctness
Ask
Whether an answer stands on cited source, points at the right region, and declines when the world cannot support it.
grounded answer precision
citation precision
abstention correctness
stale answer prevention
Operations
What it costs and how long it takes for a change in a source to reach a world an agent is allowed to read.
median and 95th-percentile time from source change to Active World
cost per 1,000 pages
cost per changed semantic unit
What a result has to carry
Every figure published here binds to a record, and five things have to be true of it before the figure may be read as a result.
A configuration nobody could edit mid-run
The model, its revision, the input mode, the prompt and the run configuration are recorded as digests taken before the run started.
A denominator
The record names the population the run covered, and so does every individual metric on it.
The raw predictions
The corpus and the run's own output are bound by digest, so the scoring can be repeated by someone who does not trust ours.
A price snapshot
Cost per page means nothing without the prices it was computed at, on the date it was computed.
The failures the run produced
Published with the run, not summarised out of it. A table that survives only because its worst row was left out is worth less than no table.
The complete receipt schema
A record missing a digest, or missing the population a rate was measured over, is refused by the build rather than rendered with a blank cell. These are the fields the validator checks.
Dataset
Which corpus was read.
Dataset version
Which cut of it, because corpora are edited.
Corpus digestsha256
sha256 over the corpus, so “the same dataset” is a checkable statement rather than a claim.
Denominator
The population every rate in the record was measured over, and how many members it has.
Model id
The exact model, never the vendor or the family name.
Model revision
The revision served on the day. Two runs of “the same model” are not the same run.
Input mode
What the model was actually given: the page image, extracted text, or both.
Prompt digestsha256
sha256 of the prompt, frozen before the run started.
Config digestsha256
sha256 of the run configuration: schema, decoding, retries, thresholds.
Compiler version
The TAVONEL build that compiled the world the metrics were read from.
Hardware
The machine. Latency and cost are properties of a machine, not of a model.
GPU
The accelerator, or “none” for a CPU run. The field is answered, never left blank.
Runtime
The serving runtime and its version.
Price snapshot
The prices in force on the day, because every cost figure expires.
Raw result digestsha256
sha256 of the unaggregated predictions, so an average can be recomputed by someone else.
World digestsha256
sha256 of the compiled world the metrics were read from.
Run receipt digestsha256
sha256 of the receipt that ties all of the above to one execution.
Date
The day the run happened, as YYYY-MM-DD.
Metrics
Every figure, with its family, its unit, and its own denominator.
Comparison basis
same_condition for a run we executed; quoted for someone else’s published score, which stays theirs.
Published failures
What the run got wrong, published with it.
Qualification rules
Freeze the configuration
Model id, revision, input mode, prompt and run configuration are pinned before the run and recorded as digests. A run whose prompt was edited while it was running is one run of two configurations, and is comparable to neither.
Publish the denominator
The record names the population it was measured over, and so does every individual metric. The same percentage over a whole corpus and over the subset that failed are two different findings that look identical without it.
Publish what failed
The weaknesses ship with the run. A hypothesis that did not hold is a finding, and a table that only survives because its worst row was left out is worth less than no table.
Reproduce before comparing
Someone else's published score is recorded as theirs, with the source, and is never restated as something we measured. A comparison waits for a run executed here under the frozen configuration above.
North Star metric
Definition
Verified Fresh Knowledge Coverage
The share of knowledge in an Active World that carries source evidence, passes validation, and agrees with the latest revision of the source it came from.
Supporting metrics
Evidence Coverage
Stale Knowledge Rate
Recompile Avoidance
Equivalence Pass Rate
Identity Review Rate
Conflict Rate
95th-percentile time from source change to Active World