Compiled context did not beat the PDF.
GDP.pdf is a public set of 100 held-out professional tasks over ten domains, published by Surge. We ran two arms of it: one frontier model answering from the original PDF, and the same model answering from the same PDF with TAVONEL’s compiled context added. Same model, same surface, same rubric, same grader. This page is the result and the conditions it was produced under.
Not comparable with published GDP.pdf scores. This run deviates from the sealed official condition in 4 ways, listed below. Nothing here may be quoted beside a leaderboard row, and no figure on this page is a reproduction of anyone else’s.
The conditions this was run under
Each row is a difference between the run manifest and the sealed protocol in eval/gdp-pdf/protocol.json.
| Condition | Official | This run |
|---|---|---|
| Surface | No tools. The model receives the question and the PDF and answers in one turn. | claude-code-read-tool — the Read tool is enabled, so the model pages through the PDF agentically.The official condition disables tools; here the Read tool is enabled so the model can page through the PDF. Judge is Opus, not gemini-3.5-flash. |
| Judge | google/gemini-3.5-flash | claude-opus-5 (non-official; the official protocol names gemini-3.5-flash)The grading model is not the one the official protocol names, so a score here and a score there were not marked by the same marker. |
| Dataset revision | 400e411fc344b1b8dd2a51e70a7ecdf469c05b3c | 8d1efb32cb57baec2265bb84da03b30654761373Sealed revision 400e411fc344b1b8dd2a51e70a7ecdf469c05b3c resolves (HTTP 302, x-linked-etag=2ba18b4f...) but its data.parquet blob is hosted on cdn-lfs-us-1.hf.co, which returns 403 AccessDenied for the signed URL via requests, huggingface_hub and curl alike. PDFs at the same revision redirect to us.aws.cdn.hf.co and download fine; only the parquet blob is stranded. Run pinned to head 8d1efb32... instead. This run is NOT official-comparable for that reason as well as the surface and judge deviations. |
| Epochs | 5 epochs per task | 1 epoch per taskA single epoch measures one draw from a non-deterministic model. The official condition averages five. |
What the run found
Adding TAVONEL's compiled context to the PDF lowered the all-criteria pass rate by 4.0 percentage points over the 100 tasks run in both arms. The compiled arm scored worse than the PDF alone, and that is the result.
Macro is the share of tasks where every rubric criterion passed — the official headline metric, and a hard one: one missed criterion out of eighteen fails the task. Micro is the mean fraction of criteria passed, which moves for partial credit. A subject, adapter or judge failure scores zero and stays in the denominator.
| Arm | n | Macro all-criteria pass | Micro mean criteria | Criteria passed |
|---|---|---|---|---|
| PDF alone | 100 | 31.0% (31 of 100) | 79.3% | 1032 of 1275 |
| PDF + compiled context | 100 | 27.0% (27 of 100) | 71.3% | 948 of 1275 |
| Delta (compiled − PDF alone) | 100 | −4.0 pp | −8.0 pp | over the tasks run in both arms |
A second denominator, not a second headline. 89 of the 100 tasks had a compiled packet that could actually be assembled. Over those 89 alone the macro delta is −1.1 pp and the micro delta +0.9 pp. That cut separates “no packet existed” from “the packet did not help”.
Per domain
Counts, not just rates. Most of these denominators are one or two tasks, which is too few to read a domain effect from — they are published because leaving them out would hide where the run is thin.
| Domain | n | PDF alone | PDF + compiled context |
|---|---|---|---|
| Construction | 10 | 3 of 10 (73.7% micro) | 3 of 10 (72.4% micro) |
| Engineering | 10 | 1 of 10 (76.8% micro) | 1 of 10 (71.8% micro) |
| Finance/Investing | 9 | 1 of 9 (78.9% micro) | 3 of 9 (75.3% micro) |
| HR | 9 | 6 of 9 (92.1% micro) | 5 of 9 (89.9% micro) |
| Healthcare | 11 | 4 of 11 (75.8% micro) | 2 of 11 (82.6% micro) |
| Insurance | 9 | 4 of 9 (86.7% micro) | 2 of 9 (51.7% micro) |
| Legal | 7 | 3 of 7 (86.2% micro) | 2 of 7 (72.1% micro) |
| Manufacturing/Supply Chains | 10 | 2 of 10 (70.8% micro) | 4 of 10 (69.4% micro) |
| Real Estate | 10 | 1 of 10 (71.6% micro) | 0 of 10 (59.5% micro) |
| STEM/Research | 15 | 6 of 15 (82.9% micro) | 5 of 15 (68.8% micro) |
Tokens, latency and cost
The money column is a list-price equivalent, never an invoice amount. The run executed on a subscription surface where the marginal cost of a call is $0; the figure is what the same tokens would have cost at the API list price in the run’s price snapshot. That snapshot carries no cache-token price, so where cost is marked incomplete the figure covers uncached input and output only and understates the run. Latency is wall-clock per task, including the model paging through the PDF.
| Arm | n | Input tokens | Output tokens | p50 | p95 | Subject (list-equiv) | Judge (list-equiv) | Cost complete |
|---|---|---|---|---|---|---|---|---|
| PDF alone | 100 | 27,229,987 (reported for 99 of 100) | 1,049,711 | 98.7s | 299.8s | $26.2471 | $3.3430 | no |
| PDF + compiled context | 100 | 24,154,317 (reported for 88 of 100) | 1,052,221 | 116.5s | 345.3s | $26.3100 | $3.2732 | no |
Hardest documents
Task ids only. The documents themselves are third-party material in a restricted run store, and no title, question or answer is published here.
| Task | Domain | Pages | Scanned | Criteria | PDF alone | PDF + compiled |
|---|---|---|---|---|---|---|
758bb8e1 | Manufacturing/Supply Chains | 5 | no | 26 | 0 of 26 (not all) | 0 of 26 (not all) |
973081e7 | Real Estate | 1 | no | 4 | 1 of 4 (not all) | 0 of 4 (not all) |
e88c1006 | Finance/Investing | 168 | no | 6 | 1 of 6 (not all) | 1 of 6 (not all) |
0a9aff44 | Healthcare | 61 | no | 10 | 2 of 10 (not all) | 2 of 10 (not all) |
31fddf6a | Construction | 19 | no | 5 | 1 of 5 (not all) | 1 of 5 (not all) |
6070d5f8 | Real Estate | 60 | yes | 5 | 3 of 5 (not all) | 0 of 5 (not all) |
cf79b3db | Insurance | 5 | yes | 8 | 5 of 8 (not all) | 0 of 8 (not all) |
bffc3940 | Engineering | 5 | no | 17 | 2 of 17 (not all) | 9 of 17 (not all) |
3c825f54 | Legal | 7 | yes | 6 | 4 of 6 (not all) | 0 of 6 (not all) |
efb22c1d | Insurance | 11 | yes | 26 | 19 of 26 (not all) | 0 of 26 (not all) |
f3537e1f | Manufacturing/Supply Chains | 40 | yes | 21 | 16 of 21 (not all) | 0 of 21 (not all) |
685a9467 | Finance/Investing | 16 | yes | 9 | 7 of 9 (not all) | 0 of 9 (not all) |
f8e8f09e | STEM/Research | 9 | yes | 14 | 11 of 14 (not all) | 0 of 14 (not all) |
678d01ed | Construction | 16 | no | 5 | 2 of 5 (not all) | 2 of 5 (not all) |
aff4bf86 | Engineering | 1 | yes | 7 | 6 of 7 (not all) | 0 of 7 (not all) |
What this does and does not show
- It shows that on 100 tasks of this set, under this surface and this grader, the compiled-context arm did not beat the model reading the PDF directly. That is a finding about our compiled context.
- It does not show a GDP.pdf score. 4 conditions differ from the official protocol, any one of which is enough to make a comparison with a published row meaningless.
- It does not show a stable effect size. 100 tasks at one epoch each, with no confidence interval, is an early read. Every rate on this page is over the tasks actually run.
- It does not show anything about the other two sealed arms. Fixed-control retrieval and the adaptive router were not run in this lane. They are not run, which is different from zero.
- It does not show the compiled context at full size. 53 of the 89 packets reached the adapter’s character budget and were cut before the model read them, so on those tasks the arm answered from a truncated packet. That is a condition of this run, recorded here; it is not a correction to the result above.
- PDF alone: 1 of 100 tasks failed in the model call and scored zero.
- PDF alone: 1 of 100 tasks could not be graded and scored zero.
- PDF alone: the cost column is incomplete. The price snapshot carries no cache-token price, so the figure covers uncached input and output only and understates the run.
- PDF alone: input tokens were reported for 99 of 100 tasks, so the token total is a floor.
- PDF + compiled context: 11 of 100 tasks scored zero because the arm could not be assembled, not because the answer was wrong. The zeros stay in the denominator, as the sealed failure policy requires.
- PDF + compiled context: 1 of 100 tasks failed in the model call and scored zero.
- PDF + compiled context: the cost column is incomplete. The price snapshot carries no cache-token price, so the figure covers uncached input and output only and understates the run.
- PDF + compiled context: input tokens were reported for 88 of 100 tasks, so the token total is a floor.
Receipts
What this page’s figures are bound to
- Artifact
nextjs/content/benchmarks/gdp-pdf.jsonsha256:0302712c8a0570492347d604c00a8090bff13b92cd96190d3844798aa258f953- Subject model
claude-opus-5- Surface
claude-code-read-tool· tools:Read- Judge
- claude-opus-5 (non-official; the official protocol names gemini-3.5-flash)
- Dataset revision run
8d1efb32cb57baec2265bb84da03b30654761373- Sealed dataset revision
400e411fc344b1b8dd2a51e70a7ecdf469c05b3c- Task set digest
002416ec3b2b51083d74b403f8bba300ba85aca622efddb38c6b1b346b4c9381- Prompt digest · compiled_context_pdf
fa7ff751335f658511c6b96d3544be41658e83c4449bb5c39abe289b2aea7840- Prompt digest · judge
b6ceadd9df489b880bcfc62a5a1cbd5761503527b2a506a9b0d6de63f0e150de- Prompt digest · native_pdf
5f78f2b58fb2441a9b067bf8ce395c204c81deddf50d408a92398db127a566ec- Epochs · seed
- 1 ·
gdp-pdf-20260921 - Cells recorded
- 200 · 100 of 100 catalogue tasks, both arms
The protocol these arms are sealed against, the source pins and the rights position on the corpus are in nextjs/eval/gdp-pdf/README.md and nextjs/eval/gdp-pdf/protocol.json in this repository, which is also what the conditions table above is diffed against. Prompts, PDFs, answers and judge rationales stay in the restricted run store; only their digests appear here.
The corpus is GDP.pdf, by Surge, used under the dataset’s own terms: not redistributed, not used for training. The reference harness is recorded as evidence and its code is not copied.
Also in the trust case: Benchmark protocol · Evidence · Reproducibility · Research notes