Yardstick Labrigorous software benchmarking
Document AI benchmark

Document AI rankings

This is Yardstick Lab's first live category. The current ranking covers 72 public documents. It scores five specialized extraction platforms alongside three frontier LLMs called directly with the same schema — the do-it-yourself baseline a team would build in-house. One row each, at its best-scoring config.

Overall structured extraction score

Macro-average field accuracy from the current ranked dataset.

One row per system, at its best-scoring config. Five specialized platforms are ranked alongside three frontier LLMs called directly with the same schema — the do-it-yourself baseline, each given its own maximum output budget. Every complete config we scored is in the table below.
72documents in the ranked dataset
8systems ranked, one row each
10configurations scored in full
1dated run behind this page
Behind the rows

Full list

The leaderboard shows one config per system. This is the full set, including each product's standard and premium settings side by side.

SystemConfigScore
DocuPipeHigh97.02%
DocuPipeStandard96.14%
Claude Sonnet 5Direct LLM91.73%
ReductoDeep Extract89.38%
ReductoStandard81.11%
ExtendDefault80.28%
GPT-5.5Direct LLM76.48%
Gemini 3.5 FlashDirect LLM72.98%
Pulse AIpulse-ultra-2 + extended reasoning70.95%
UnstructuredAuto partitioning67.67%
Slice rankings

Where the score separates

These slices come from the same current run and show evidence for specific buyer workflows.