Automate document processing with AI
Skilled people spend their days copying amounts, dates and references from a PDF into software. Current models read these documents, extract structured data and flag what they are unsure about. The project is less about extraction than about control: knowing when to trust the machine.
The problem
OCR and template-based recognition work as long as documents look alike. As soon as suppliers, formats or languages vary, maintaining templates costs more than typing.
Multimodal language models read a document the way a person would, with no template. But they can be confidently wrong. Reliable processing therefore combines extraction, business validation rules and targeted human review.
How we work
-
1
Measure the current state
Volume per document type, processing time, current error rate, cost of an error. These figures are yours, and they decide what is worth automating and in which order.
-
2
Build a reference sample
A few hundred real documents with the correct values. On this sample we compare approaches and measure, field by field, what extraction is really worth.
-
3
Extract, then check
Extraction into a strict schema, followed by business checks: consistent totals, plausible VAT, known supplier, existing purchase order. A document that passes every check is processed untouched; the others go to a review queue with the doubtful field highlighted.
-
4
Integrate and monitor
Data is sent to your ERP, accounting software or document management system through APIs. We track the straight-through rate and human corrections, which feed back into the system.
What makes these projects fail
Aiming for zero touch
Trying to automate everything means accepting silent errors. The reasonable goal is to process simple cases untouched and make hard cases faster to validate.
Scoring per document, not per field
A document that is “95% correct” with a wrong IBAN is a wrong document. Accuracy is measured field by field, weighted by the cost of the error.
Neglecting traceability
For each value you must be able to find the source document, the extracted value, the check applied and the person who validated it. That is what an auditor will ask for.
What you receive
- Baseline measurement and annotated reference sample
- Extraction pipeline and business validation rules
- Review interface for doubtful cases
- Connectors to your ERP, accounting or document management system
- Indicators: straight-through rate, corrections, lead times
Frequently asked questions
Which document types can be processed?
Invoices, credit notes, purchase and delivery orders, contracts, identity documents, forms, statements, letters. Handwritten or very poor quality documents remain harder: we check on your samples before committing.
Do we need a model trained on our documents?
Rarely at first. Current general-purpose models give good results without training. Fine-tuning can make sense later for a very specific, high-volume document type, once the reference sample exists.
What happens to the documents sent to the model?
We select providers that commit contractually not to reuse data, with European hosting when required, or a model deployed on your infrastructure for the most sensitive documents.
Is your situation close to this one?
Describe it in a few lines. We will tell you whether AI is the right answer — and we will also tell you when it is not.
Talk about your project