Compiled context did not beat the PDF.

GDP.pdf is a public set of 100 held-out professional tasks over ten domains, published by Surge. We ran two arms of it: one frontier model answering from the original PDF, and the same model answering from the same PDF with TAVONEL’s compiled context added. Same model, same surface, same rubric, same grader. This page is the result and the conditions it was produced under.

Not comparable with published GDP.pdf scores. This run deviates from the sealed official condition in 4 ways, listed below. Nothing here may be quoted beside a leaderboard row, and no figure on this page is a reproduction of anyone else’s.

The conditions this was run under

Each row is a difference between the run manifest and the sealed protocol in eval/gdp-pdf/protocol.json.

Deviations from the official condition
ConditionOfficialThis run
SurfaceNo tools. The model receives the question and the PDF and answers in one turn.claude-code-read-tool — the Read tool is enabled, so the model pages through the PDF agentically.The official condition disables tools; here the Read tool is enabled so the model can page through the PDF. Judge is Opus, not gemini-3.5-flash.
Judgegoogle/gemini-3.5-flashclaude-opus-5 (non-official; the official protocol names gemini-3.5-flash)The grading model is not the one the official protocol names, so a score here and a score there were not marked by the same marker.
Dataset revision400e411fc344b1b8dd2a51e70a7ecdf469c05b3c8d1efb32cb57baec2265bb84da03b30654761373Sealed revision 400e411fc344b1b8dd2a51e70a7ecdf469c05b3c resolves (HTTP 302, x-linked-etag=2ba18b4f...) but its data.parquet blob is hosted on cdn-lfs-us-1.hf.co, which returns 403 AccessDenied for the signed URL via requests, huggingface_hub and curl alike. PDFs at the same revision redirect to us.aws.cdn.hf.co and download fine; only the parquet blob is stranded. Run pinned to head 8d1efb32... instead. This run is NOT official-comparable for that reason as well as the surface and judge deviations.
Epochs5 epochs per task1 epoch per taskA single epoch measures one draw from a non-deterministic model. The official condition averages five.

What the run found

Adding TAVONEL's compiled context to the PDF lowered the all-criteria pass rate by 4.0 percentage points over the 100 tasks run in both arms. The compiled arm scored worse than the PDF alone, and that is the result.

Macro is the share of tasks where every rubric criterion passed — the official headline metric, and a hard one: one missed criterion out of eighteen fails the task. Micro is the mean fraction of criteria passed, which moves for partial credit. A subject, adapter or judge failure scores zero and stays in the denominator.

Headline, per arm
ArmnMacro all-criteria passMicro mean criteriaCriteria passed
PDF alone10031.0% (31 of 100)79.3%1032 of 1275
PDF + compiled context10027.0% (27 of 100)71.3%948 of 1275
Delta (compiled − PDF alone)100−4.0 pp−8.0 ppover the tasks run in both arms

A second denominator, not a second headline. 89 of the 100 tasks had a compiled packet that could actually be assembled. Over those 89 alone the macro delta is −1.1 pp and the micro delta +0.9 pp. That cut separates “no packet existed” from “the packet did not help”.

Per domain

Counts, not just rates. Most of these denominators are one or two tasks, which is too few to read a domain effect from — they are published because leaving them out would hide where the run is thin.

Macro all-criteria pass, per domain and arm
DomainnPDF alonePDF + compiled context
Construction103 of 10 (73.7% micro)3 of 10 (72.4% micro)
Engineering101 of 10 (76.8% micro)1 of 10 (71.8% micro)
Finance/Investing91 of 9 (78.9% micro)3 of 9 (75.3% micro)
HR96 of 9 (92.1% micro)5 of 9 (89.9% micro)
Healthcare114 of 11 (75.8% micro)2 of 11 (82.6% micro)
Insurance94 of 9 (86.7% micro)2 of 9 (51.7% micro)
Legal73 of 7 (86.2% micro)2 of 7 (72.1% micro)
Manufacturing/Supply Chains102 of 10 (70.8% micro)4 of 10 (69.4% micro)
Real Estate101 of 10 (71.6% micro)0 of 10 (59.5% micro)
STEM/Research156 of 15 (82.9% micro)5 of 15 (68.8% micro)

Tokens, latency and cost

The money column is a list-price equivalent, never an invoice amount. The run executed on a subscription surface where the marginal cost of a call is $0; the figure is what the same tokens would have cost at the API list price in the run’s price snapshot. That snapshot carries no cache-token price, so where cost is marked incomplete the figure covers uncached input and output only and understates the run. Latency is wall-clock per task, including the model paging through the PDF.

Tokens, latency and list-price-equivalent cost, per arm
ArmnInput tokensOutput tokensp50p95Subject (list-equiv)Judge (list-equiv)Cost complete
PDF alone10027,229,987 (reported for 99 of 100)1,049,71198.7s299.8s$26.2471$3.3430no
PDF + compiled context10024,154,317 (reported for 88 of 100)1,052,221116.5s345.3s$26.3100$3.2732no

Hardest documents

Task ids only. The documents themselves are third-party material in a restricted run store, and no title, question or answer is published here.

Ranked by the PDF-alone arm’s mean criteria
TaskDomainPagesScannedCriteriaPDF alonePDF + compiled
758bb8e1Manufacturing/Supply Chains5no260 of 26 (not all)0 of 26 (not all)
973081e7Real Estate1no41 of 4 (not all)0 of 4 (not all)
e88c1006Finance/Investing168no61 of 6 (not all)1 of 6 (not all)
0a9aff44Healthcare61no102 of 10 (not all)2 of 10 (not all)
31fddf6aConstruction19no51 of 5 (not all)1 of 5 (not all)
6070d5f8Real Estate60yes53 of 5 (not all)0 of 5 (not all)
cf79b3dbInsurance5yes85 of 8 (not all)0 of 8 (not all)
bffc3940Engineering5no172 of 17 (not all)9 of 17 (not all)
3c825f54Legal7yes64 of 6 (not all)0 of 6 (not all)
efb22c1dInsurance11yes2619 of 26 (not all)0 of 26 (not all)
f3537e1fManufacturing/Supply Chains40yes2116 of 21 (not all)0 of 21 (not all)
685a9467Finance/Investing16yes97 of 9 (not all)0 of 9 (not all)
f8e8f09eSTEM/Research9yes1411 of 14 (not all)0 of 14 (not all)
678d01edConstruction16no52 of 5 (not all)2 of 5 (not all)
aff4bf86Engineering1yes76 of 7 (not all)0 of 7 (not all)

What this does and does not show

  • It shows that on 100 tasks of this set, under this surface and this grader, the compiled-context arm did not beat the model reading the PDF directly. That is a finding about our compiled context.
  • It does not show a GDP.pdf score. 4 conditions differ from the official protocol, any one of which is enough to make a comparison with a published row meaningless.
  • It does not show a stable effect size. 100 tasks at one epoch each, with no confidence interval, is an early read. Every rate on this page is over the tasks actually run.
  • It does not show anything about the other two sealed arms. Fixed-control retrieval and the adaptive router were not run in this lane. They are not run, which is different from zero.
  • It does not show the compiled context at full size. 53 of the 89 packets reached the adapter’s character budget and were cut before the model read them, so on those tasks the arm answered from a truncated packet. That is a condition of this run, recorded here; it is not a correction to the result above.
  • PDF alone: 1 of 100 tasks failed in the model call and scored zero.
  • PDF alone: 1 of 100 tasks could not be graded and scored zero.
  • PDF alone: the cost column is incomplete. The price snapshot carries no cache-token price, so the figure covers uncached input and output only and understates the run.
  • PDF alone: input tokens were reported for 99 of 100 tasks, so the token total is a floor.
  • PDF + compiled context: 11 of 100 tasks scored zero because the arm could not be assembled, not because the answer was wrong. The zeros stay in the denominator, as the sealed failure policy requires.
  • PDF + compiled context: 1 of 100 tasks failed in the model call and scored zero.
  • PDF + compiled context: the cost column is incomplete. The price snapshot carries no cache-token price, so the figure covers uncached input and output only and understates the run.
  • PDF + compiled context: input tokens were reported for 88 of 100 tasks, so the token total is a floor.

Receipts

What this page’s figures are bound to
Artifact
nextjs/content/benchmarks/gdp-pdf.json
sha256:0302712c8a0570492347d604c00a8090bff13b92cd96190d3844798aa258f953
Subject model
claude-opus-5
Surface
claude-code-read-tool · tools: Read
Judge
claude-opus-5 (non-official; the official protocol names gemini-3.5-flash)
Dataset revision run
8d1efb32cb57baec2265bb84da03b30654761373
Sealed dataset revision
400e411fc344b1b8dd2a51e70a7ecdf469c05b3c
Task set digest
002416ec3b2b51083d74b403f8bba300ba85aca622efddb38c6b1b346b4c9381
Prompt digest · compiled_context_pdf
fa7ff751335f658511c6b96d3544be41658e83c4449bb5c39abe289b2aea7840
Prompt digest · judge
b6ceadd9df489b880bcfc62a5a1cbd5761503527b2a506a9b0d6de63f0e150de
Prompt digest · native_pdf
5f78f2b58fb2441a9b067bf8ce395c204c81deddf50d408a92398db127a566ec
Epochs · seed
1 · gdp-pdf-20260921
Cells recorded
200 · 100 of 100 catalogue tasks, both arms

The protocol these arms are sealed against, the source pins and the rights position on the corpus are in nextjs/eval/gdp-pdf/README.md and nextjs/eval/gdp-pdf/protocol.json in this repository, which is also what the conditions table above is diffed against. Prompts, PDFs, answers and judge rationales stay in the restricted run store; only their digests appear here.

The corpus is GDP.pdf, by Surge, used under the dataset’s own terms: not redistributed, not used for training. The reference harness is recorded as evidence and its code is not copied.

Also in the trust case: Benchmark protocol · Evidence · Reproducibility · Research notes