RECORDEVIDENCE & BOUNDARY

What we enforce,
and what we actually measured.

Everything on the front page is a demonstration. Nothing here is. There are no customer logos anywhere on this site and no certifications — a brand rule bars them without registered evidence — so this is our own record instead, including the part of it that did not work.

WHAT WE MEASURED, AND WHAT WE DID NOT

MEASURED

Recovery changes the outcome

On a public benchmark with an unmodified scoring path, the recovery runtime moved a document extraction score substantially. Our own measurement, published with its confidence interval, and never placed beside a competitor's number as if reproduced.

MEASURED

Compilation refuses more than it emits, sometimes

Of a thousand documents offered in one campaign, four hundred and four were refused, every one for a link the compiler could not resolve. A vault with a broken link is not emitted, by design.

NOT SUPPORTED

Blind quality detection failed

We tested whether prediction-only signals could pick the worst documents without ground truth. They could not beat ranking by length alone. Published as unsupported, and not shipped as a feature.

BUILT, NOT PROVEN

Most thresholds are uncalibrated

Tests show the code does what its author intended. They do not show a threshold is right. Nothing here presents an uncalibrated threshold as a measured result.

Two of the four entries above are things that did not work or are not proven. That ratio is the point. A record that only listed the wins would be a claim about our marketing rather than about our engineering, and it would tell you nothing you could check.