Intelligent document processing UK buyers have a genuinely better market than they did two years ago: commercial extraction products now handle invoices, purchase orders and standard forms well enough that building your own would be wasteful. The decision has therefore moved. It is no longer whether the technology works, it is whether your documents look like the ones the products were built for. For expert-led businesses, they often do not.

What off-the-shelf products do well

High-volume, standardised document types with a common structure and a well-understood set of fields. Invoices are the canonical case: millions of examples, a small stable field set, and a clear right answer. The same applies to purchase orders, delivery notes, standard application forms and identity documents. If your process is dominated by documents like these, buy. Configuration will take days, accuracy will be high, and the ongoing cost is a subscription rather than a maintenance liability.

Modern products also handle moderate variation in layout well, so the old objection that every supplier's invoice looks different has largely dissolved.

Where they struggle, and why it matters to specialists

Four situations reliably defeat generic tools. Documents where meaning depends on domain context: a test certificate where whether a result passes depends on which standard applies and which revision of it. Documents where the extraction is really an assessment: pulling the effective obligations out of a contract or a consent, which is judgement wearing the costume of data entry. Long technical documents where the field you want appears in different sections with different phrasing, so position offers no help. And documents where the required output is not a set of fields at all but a comparison, such as whether this revision differs materially from the last.

Expert-led businesses run on exactly these document types: specifications, standards, test reports, survey records, regulatory submissions and consents. That is why the generic evaluation so often ends at a promising demo and a disappointing pilot, a pattern we have written about in why generic platforms fall short.

The honest build-or-buy test

Run the products against fifty of your real documents, deliberately including the awkward ones, and score field-level accuracy yourself rather than accepting a vendor summary. Three outcomes. Above roughly ninety-five per cent on the fields that matter, buy and configure. Between eighty and ninety-five, look at what fails: if it is a handful of fields, a small custom layer over the product often beats a full build. Below eighty on your core fields, and the gap is usually domain understanding rather than tuning, which is the case for building.

Weigh the ongoing cost honestly on both sides. A subscription priced per page can become the dominant cost at volume, while a custom system trades that for a maintenance obligation that never quite disappears. Neither is free.

What a custom build looks like

Less exotic than it sounds. Commercial models do the reading; the custom work is the domain layer around them: your document taxonomy, your validation rules, the standards and thresholds that decide whether an extracted value is acceptable, and a review interface where a specialist corrects the small proportion of cases the system is unsure about. Those corrections should feed back so the system improves rather than repeating the same misreads. Our property document AI work followed that shape, and the differentiator was the domain rules rather than the extraction.

Where the process is repeatable and the bottleneck is document handling rather than judgement, that is our AI Workflow Automation work: the extraction, the checks and the exception routing built into the process you already run, without a platform migration. The wider view of what we build for document-heavy operations sits on our document-heavy operations page.

Frequently asked questions

How accurate is intelligent document processing?

On standard forms and invoices, commercial products commonly exceed ninety-five per cent field accuracy. On specialist technical documents the range is far wider and depends on how much domain context the field requires. The only number worth trusting is the one you measure on your own documents.

Should we build or buy document processing?

Test before deciding. Run candidate products on fifty real documents including the difficult ones and score the fields that matter. High accuracy means buy. Persistent failures concentrated in domain-dependent fields mean the gap is understanding rather than tuning, which is the case for building.

Can AI check documents as well as extract from them?

Yes, and checking is often the larger prize. Comparing extracted values against limits, standards or previous revisions catches the errors that cost money, and it has a verifiable right answer. Businesses frequently underestimate checking because extraction is the more familiar pitch.

What about confidential client documents?

Ask any supplier where documents are processed, how long they are retained, and whether their contents may be used to improve a shared model. For sensitive or personal material a system that processes within your own infrastructure removes the question entirely, and that requirement alone sometimes decides build over buy.

Not sure where AI fits in your business?

The AI Opportunity Finder maps your highest-value starting point in a few minutes, with no sales call required.

Find your best starting point