Guides · Document processing

How accurate is document extraction, really?

Short answer

Any accuracy number quoted without three things attached — which field, which document mix, and what counts as correct — is marketing. The figure that actually predicts your outcome is the straight-through rate: the share of documents that pass end to end without a person touching them.

Per field
Accuracy differs enormously between fields on one page
Your documents
The only sample whose result predicts yours
Straight-through
The number that turns into hours saved
By Max Pochinsky · Updated August 2026 · Written for people scoping a project, not for search engines

Why one number cannot describe it

Take an invoice. The total is a single well-formatted figure in a predictable place and is read almost perfectly. The supplier name has to be matched against your own list of suppliers, where two entries differ by a legal suffix. The line items sit in a table that spans a page break, and the delivery date might be handwritten in a margin. These four fields do not have one accuracy — they have four, and they are not close to each other.

Then there is the question of what correct means. Is a supplier matched to the right parent company but the wrong subsidiary correct? Is a date read correctly but in the wrong format correct? Two vendors can quote very different numbers on the same documents purely because they answered that question differently, and neither is lying.

The measurements that are worth having

01

Field-level accuracy, per field. How often each individual field is exactly right. This is what tells you where the work is, and it is the only view that lets you improve anything.

02

Document-level accuracy. How often every field on a document is right at once. Always lower than the field average, and much lower when there are many fields — this is the number that surprises people.

03

Straight-through rate. The share of documents that complete the whole process without human intervention. This is what converts into saved hours, and it is the number to put in a business case.

04

Escaped error rate. How often something wrong is accepted with high confidence and reaches your accounts. Rare, and by far the most expensive thing to get wrong. Measure it separately.

05

Confidence calibration. When the system says it is sure, is it? A well-calibrated system with modest accuracy is more useful than a confident one with high accuracy, because you can trust it to ask.

What actually moves the numbers

Input quality dominates. A native PDF from an accounting system reads close to perfectly. A photograph of a crumpled receipt taken at an angle in poor light does not, and no model fixes the information that was never captured. If the process can be changed so documents arrive digitally, that single change usually beats any amount of tuning.

After that: how many layouts you receive, how much of the content is handwritten, whether tables cross page boundaries, how many languages are in the mix, and — the underrated one — how good your reference data is. Half of what looks like extraction error is really matching error: the field was read correctly and then matched against a supplier list with three spellings of the same company.

How to establish the real number

Take a hundred to two hundred real documents, sampled to reflect your actual mix rather than the tidy ones. Have someone key them by hand into the fields you care about — this is your ground truth, and it is the part people try to skip. Then run the system against the same set and compare field by field.

Two days of work produces a number you can act on, and it almost always changes the plan. It shows which fields are safe to automate today, which need a confidence threshold, and which should stay manual for now. It also surfaces the disagreements between your own reviewers, which is a finding in itself — when two experienced people key a field differently, no system is going to score well against either of them.

We run this measurement during the free mini audit, on your documents, before quoting anything. Any accuracy figure we put in a proposal comes from that run.

Designing for the errors you will have

Perfect extraction is not the goal and not available. A usable system is one where the errors are cheap: a confidence threshold below which a document goes to a person, a review screen that shows the source image beside the extracted value so checking takes seconds, and a rule that anything above a certain amount is verified regardless of confidence.

Built that way, an eighty per cent straight-through rate is a very good outcome — four in five documents never touched, the remaining fifth reviewed quickly with the answer already filled in. Chasing the last few per cent usually costs more than the manual handling of those documents ever did.

Follow-up questions

What people ask next.

Not before seeing your documents, and neither can anyone else honestly. On clean native PDFs with a stable layout, key fields do reach that kind of level. On mixed scans, photographs and handwriting, they do not. We measure first and put the measured figure in the quote.

Still unsure whether your process is worth automating?

Bring us the process. We take it apart with you at no charge and give you a straight answer, including when the answer is no.