Key takeaways
- A system can extract most fields correctly and still send many documents to a person, because one wrong field is enough to stop a document.
- In a 2025 test on 102 invoices, one pipeline reached 94% overall field accuracy and another 63%, yet 93% and 80% of their invoices passed an arithmetic consistency check.
- We found no openly published, primary benchmark of straight-through rates for invoices or claims, and the vendor figures in circulation conflict.
- Human review is itself error-prone, so the exception queue needs design, sampling and accountability rather than a promise that someone will check.
In 2025, researchers published a comparison of two ways to extract data from invoices on a set of 102 invoices in English and German layouts, some digitally born and some scanned. One pipeline, built on Docling, reached 63% overall field accuracy. A second, LlamaExtractor, reached 94%.1 The gap between 63% and 94% is the kind of number that ends up on a slide. It is not the number that decides what a finance or claims operation pays.
That number is the share of documents that go from inbox to ledger with no person touching them, known as straight-through processing, or STP. A document with one wrong field out of many still needs a human, and the human costs the same whether one field or ten were wrong. For an operations leader, the question to put to any vendor or internal team is therefore how many documents pass with no touch, and what happens to the rest.
Why field accuracy flatters the system
Field accuracy is counted per field. A document is only straight-through if every field it needs is right, or if the wrong ones are caught. The two measures diverge as soon as a document has more than a few fields, because errors in different fields add up document by document.
Take the simplest model. If each field were right 98% of the time and fields failed independently, a document with 10 fields would be fully right about 82% of the time. Real fields are not independent, since a poor scan degrades several at once, so this is an illustration of the shape of the problem rather than a prediction. The point survives the simplification: a system that looks excellent field by field can still leave nearly one document in five for a person.
The 2025 invoice study shows the same pattern from the other direction. Its authors also ran a consistency check on each invoice, testing whether net amount, tax and rounding add up to the gross amount. LlamaExtractor passed that check on 93% of invoices, and the Docling pipeline on 80%, with 20% failing.1 The check is not a measured STP rate, and it should not be quoted as one. It does show that a document-level test and a field-level score answer different questions.
Figure 1
Field accuracy versus a document-level check, 102 invoices
| LlamaExtractor | Docling pipeline | |
|---|---|---|
| Overall field accuracy | 94% | 63% |
| Line-item accuracy | 91% | 58% |
| Invoices passing the net + tax + rounding = gross check | 93% | 80% |
Look at the Docling row. Overall field accuracy was 63%, and 80% of its invoices still passed the arithmetic check.1 A passing check, in other words, is weaker evidence than it appears, because amounts can add up and still be attached to the wrong line or the wrong vendor. Any rule that lets a document through on a consistency test alone should be treated as a gate to be audited, not as proof.
How thresholds and exception queues work
Most extraction systems attach a confidence score to each field. A threshold turns that score into a decision. Above it, the field is accepted. Below it, the document goes to a queue where a person looks at the page, the extracted values and the places they came from. Where the threshold sits is a business choice: raising it sends more documents to people and lowers the error rate that reaches the ledger, and lowering it does the reverse.
Two design details matter. The first is that the reviewer should see the exact spot on the page that each value was read from. The DocILE benchmark, a public dataset of 6.7k annotated business documents with 55 field classes, makes the same argument for its task, noting that localization is crucial for human-in-the-loop interactions, auditing and other processing.2 A reviewer who must hunt for the value on the page takes longer, and a rushed reviewer approves more mistakes.
The second detail is that the threshold should vary by field. A wrong invoice number is usually cheap to fix. A wrong bank account or total is not. Setting one global threshold treats them alike, so the cheap errors clog the queue while the expensive ones leak through. A better arrangement sets stricter thresholds on the fields where an error costs money and lets the cheap fields pass with a looser rule and a periodic sample.
The DocILE authors also point out that most public datasets are small and that many results are published on private datasets.2 That matters here, because it is why a buyer cannot read a straight-through rate off a leaderboard. The rate depends on your suppliers, your layouts and your scan quality, and only your own documents can measure it.

What a reviewer actually catches
The case for the exception queue assumes the person in it catches what the machine missed. The research on human error is less comforting. In a review of what is known about spreadsheet errors, Raymond Panko reports that detection and correction rates approaching 90% occur only in the simplest processes, such as proofreading spelling errors that are not valid words. Detection falls to about 70% when the misspelling is itself a valid word, and a 1984 study by Allwood found only about half of logical errors in mathematics were caught.3 That research concerns spreadsheets and proofreading, not invoice approval, so it sets a plausible range rather than measuring your queue.
A systematic review of automation bias in clinical decision support covers a different field again and points the same way. Of 13,821 papers retrieved, 74 met its inclusion criteria, and the authors note that there is often a failure to recognize the new errors that such systems introduce.4 The mitigators it discusses are training, accountability and design choices such as showing updated confidence levels on the output. Those are findings from clinical settings, and nobody has shown that the effect size transfers to accounts payable. The mechanism is easy to imagine, though: a reviewer who sees a pre-filled form and a high confidence score is inclined to agree.
What straight-through is worth
The arithmetic is straightforward, and the weak link is the straight-through rate itself. We searched for an openly published, primary benchmark of touchless or straight-through rates for invoices and claims and did not find one we could verify. Secondary articles quote rates and per-document costs that cite industry surveys, and the figures disagree with one another, so we do not use any of them. If a vendor quotes a rate, ask for the definition, the document mix and the period.
Newmind’s document-processing benchmark report sets out cost per document, straight-through rate, field accuracy and turnaround figures to compare against your intake workflows.

What the data does not say
The invoice study is small: 102 invoices, curated, in two languages. Its authors note that large industry invoice datasets are scarce because of privacy, and a result on 102 documents will not transfer cleanly to a supplier base of thousands of layouts.1 The two pipelines it compares are also two snapshots from 2025, and extraction tools change quickly.
A high straight-through rate can also be a failure. If the threshold is set loosely, the rate rises because mistakes are being approved, and the cost reappears later as duplicate payments, disputes and corrections. Report straight-through together with the error rate found in a sample of the documents that passed untouched. One number without the other can be gamed, usually without anyone intending it.

What to do next
- Define straight-through in writing: a document counts only if no person touched it and a later sample finds it correct.
- Measure the baseline on a month of your own documents, split by supplier or form type, before you test any tool.
- Set confidence thresholds field by field, strictest where an error moves money, and review the queue size each week.
- Show reviewers the source location of every value, and sample a share of auto-approved documents for audit.
- Track the exception reasons, since the same few causes usually account for most of the queue and can often be fixed upstream.
If you want a first estimate of where AI could cut cost in your own document workflows, the free pre-audit is a short questionnaire that returns one.
Newmind Partners
Find out where AI would pay in your workflows
Newmind Partners designs and builds AI workflows that cut operating cost. Start with the free pre-audit for a first estimate, or run a Feasibility audit for a scored report on one workflow.
Sources
- Yashwant, Dubey, Paikray and Thulsiram, “Invoice Information Extraction: Methods and Performance Evaluation” (arXiv 2510.15727, October 2025), curated set of 102 invoices in English and German. arxiv.org
- Šimsa et al., “DocILE Benchmark for Document Information Localization and Extraction” (arXiv 2302.05658, ICDAR 2023), 6.7k annotated business documents. arxiv.org
- Panko, “Spreadsheet Errors: What We Know. What We Think We Can Do” (EuSpRIG, July 2000; arXiv 0802.3457), review of human-error research applied to spreadsheets. arxiv.org
- Goddard, Roudsari and Wyatt, “Automation bias: a systematic review of frequency, effect mediators, and mitigators”, Journal of the American Medical Informatics Association 19(1):121-127 (2012), 74 of 13,821 retrieved papers included. academic.oup.com




