Yardstick Labrigorous software benchmarking
Document AI benchmark

Every document, every score

A ranking you cannot audit is an opinion. This is the full working of the document AI leaderboard: all 72 documents, what each system scored on each one, and an explorer where you can open any document alongside its answer key and every system's extraction to judge it for yourself. The column averages at the bottom of the table are the headline ranking numbers; no document is removed from them.

How to read this

The scoring, in one paragraph

Each document has a fixed extraction schema and a human-checked answer key. A system's score on a document is the share of fields it got right, so 100 means every field matched the key. The best score on each row is highlighted. Where a system could not process a file at all, that counts as a zero rather than being dropped from the average.

Document DocuPipeClaude Sonnet 5ReductoExtendGPT-5.5Gemini 3.5 FlashPulse AIUnstructured Explore
Average across all 72 documents
97.0291.7389.3880.2876.4872.9870.9567.67 See ranking

Every document, answer key, and system prediction is hosted here — open any of them in the data explorer. To reproduce the scores independently, download the replication kit: the answer keys, extraction schemas, and the scorer, so you can re-run it and get these numbers back.

Corpus

Where these documents come from

Every document is a real file with a traceable public source: SEC filings, utility bills, government forms, published annual reports, and vendor sample documents. Provenance and licensing for each one is recorded in the replication kit. The corpus deliberately includes spreadsheets, XML, scans, photographs, and documents in a dozen languages, because the interesting part of this category is what happens on the hard tail, not on a clean English invoice.