Yardstick Labrigorous software benchmarking
Document AI

Which document AI tool is actually most accurate?

We gave each ranked system — specialized extraction platforms and frontier LLMs called directly — the same 72 documents and schemas, scored the output field by field, and published all of it. Real invoices, bank statements, dense directories, rotated scans, handwriting, and twelve languages. If a tool could not read a file, that scored zero rather than quietly disappearing from its average.

How to use this

Start wherever your question is

The ranking is the summary. Everything under it is there so you can disagree with the summary using our own data.

The answer

Rankings

The overall table, plus the slices that matter to buyers: array-heavy documents, reconciliation, and hard languages and layouts.

See the rankings
The working

Evidence

All 72 documents, each tool's score on each one, and an explorer to open every file with its answer key and each extraction. The column averages are exactly the headline numbers.

Check every score
By workflow

Use Cases

The 72 documents grouped into the families buyers actually run — invoices, statements, payroll, healthcare, logistics, and dense registers — each scored on its own, with the leader per family.

See use cases
Open evidence

Check every number yourself.

The documents, the per-file scores, and the failures are all published right here. Open any document with its answer key and every system's extraction, or download the replication kit and re-run the scorer yourself. Nothing here is a number you have to take on faith.

Explore the data