Skip to content

Technology: OCR and data extraction

Invoices and documents read automatically, with nothing retyped.

Invoices, delivery notes and shipping statements arrive as PDFs, sometimes scanned. Someone opens them again and retypes the data. The software reads them, extracts the fields you need and checks them before they are used.

“He understood the assignment, he asked good relevant questions and delivered excellent work that did not need correcting.” Verified client, data extraction from PDFs

  • OCR
  • PDF reading
  • Structured e-invoices
  • Checks on totals

At a glance

Used for
Reading invoices, delivery notes, statements and orders without retyping
Not needed for
Structured e-invoices (XML): the data is already in fields
Where it runs
On your own servers if you like: documents stay in the company
The key point
Every extracted value is checked, never taken on trust

How it works

From PDF to ready data, in 4 steps.

Example: OCR reads the date, the code and the total from the PDF invoice. The check finds the code, confirms the total and flags the price to a person for review. Only data that adds up reaches the ERP.

The problem

The data is already there, printed in a PDF.

A supplier sends the invoice as a PDF, the courier sends the shipping statement, the warehouse receives the delivery note. The figures are all there, but to use them someone has to retype them: into a spreadsheet, into the ERP, into a check.

With a few documents it is manageable. With hundreds a month it becomes a full-time job, and every manual copy brings a few mistakes with it.

In detail

What happens at each step.

There is no single OCR that fits everything. The first step is knowing which document you are looking at.

  1. 01

    The document

    Structured e-invoice, PDF with text inside, or a scan: three different cases, three different paths. In a structured e-invoice the data is already in fields and OCR is not needed.

  2. 02

    The reading

    In a PDF with text the data is read directly; only scans need OCR. Then the fields are extracted, with one template per supplier or courier.

  3. 03

    The check

    Lines must add up to the total, tax must match, codes must exist. Whatever does not add up does not pass: it goes to a person.

  4. 04

    The data, ready

    The data goes where it is needed: the ERP, a spreadsheet, a comparison with order or contract.

When it pays off

It pays off when documents are many and look alike.

It makes sense if

  • You receive hundreds of documents a month, or a few dozen very long ones
  • Layouts repeat: the same suppliers, the same couriers
  • The data then has to be compared or loaded into an ERP
  • Today someone retypes it by hand

You do not need it if

  • You only receive structured e-invoices and your ERP already reads them
  • You get only a few documents a month
  • Every document is different and a person has to read it anyway

An example with numbers

300 invoices a month: where the time goes.

Reference figures, to redo with your own.

15 hours

a month retyping

300 invoices at 3 minutes each, before any checking.

30 invoices

to look at, afterwards

If one in ten has something that does not add up, the person opens only those.

8 lines

wrong out of 400

An OCR that is 98% right still gets 8 lines out of 400 wrong. That is why every line is checked.

The limits

What OCR does not do on its own.

It makes mistakes on scans

A 1 read as a 7, a 0 as an 8. That is why numbers are checked against totals, never taken on trust.

Every layout has to be learned

A new supplier or a changed layout has to be added. We plan for it at the start instead of finding out later.

It does not understand context

It reads “10% discount”, but whether that discount was due is in the contract, not the document.

Doubts stay with a person

When a value fails the checks, the system flags it instead of guessing.

Frequently asked

Questions on this topic.

Do structured e-invoices need OCR?

No. An XML e-invoice, like the Italian SdI format or a Peppol UBL invoice, already has its data in structured fields: it is read directly, with no OCR. OCR is for scanned PDFs and documents that arrive only as images.

How accurate is it?

On a PDF with text inside, reading is exact. On scans it depends on quality: that is why every extracted value goes through checks on totals and codes, and whatever does not add up goes to a person.

Do the documents leave the company?

Not necessarily. The software can run on your own servers without sending documents to outside services.

Is artificial intelligence needed?

Not always. For recurring layouts, templates and rules are enough. AI helps when layouts change often, but the result still has to be checked.

From how many documents does it pay off?

Usually from a few hundred a month. Sooner, if the data then has to be compared with orders, price lists or contracts.

Start with the problem

Send us a sample of your documents.

An invoice, a delivery note or a shipping statement: we will tell you whether it can be read automatically, and how.

  • The first check is free and commits you to nothing
  • Every message gets read, there is no call centre
  • Fixed price and payment by milestone
  • Italian and English, VAT invoices
  • No phone call unless you want one
  • If nothing should be built, we will say so

Prefer to write directly? [email protected]

No newsletter, no unsolicited calls. A direct reply from the developer, within one working day.