Define correct before you choose anything
Document projects that start with a product demonstration tend to end in argument, because nobody agreed what accuracy means. Is a supplier name correct if it reads Limited where the document says Ltd. Is a total correct if the currency is missing. Does a missing purchase order reference count as an extraction failure or a document failure. These sound pedantic and they are exactly the questions that decide whether a pilot is judged a success or a disappointment.
We build a labelled set first, taken from your real documents rather than samples, covering the ordinary cases and a deliberate share of the awkward ones: the poor scan, the handwritten annotation, the two page invoice, the credit note, the supplier who changed their template last year. That set becomes the benchmark, and every subsequent change is measured against it. It also makes vendor claims testable, since a general accuracy figure means nothing against your particular mix of documents.
Definition of done is a business decision as much as a technical one. For accounts payable it might be that a stated share of invoices post without human touch while the rest are cleared within an agreed time. That is a target you can measure and argue about honestly. What we avoid is a project whose success criterion is that the system is live, because that criterion is met perfectly well by a system nobody uses.
- A labelled benchmark set built from your own documents, including the awkward ones
- Field level definitions of correct, agreed with the people who do the work today
- Success stated as a measurable outcome rather than as the system being live
- Vendor and model claims tested against your document mix, not against general figures
- The benchmark retained, so later changes can be measured rather than assumed