Skip to content
Data & documents

Text recognition in practice: what it reads, what it guesses

Text recognition realistically assessed: clean sources versus carbon copies, stamps and handwriting, measuring quality, fields to extract, effort per type.

14 min read TexterkennungBelegerfassungEingangsrechnungenDokumentenmanagementDatenqualität

Few digitisation projects sound as simple as this one: scan the documents, run text recognition over them, take the data across, put the paper away. In a demonstration it looks effortless, because the samples shown are clean. In daily operation the pile also contains carbon copies, delivery notes with a stamp across the line items and handwritten notes in the margin. That is where it is decided whether capture actually saves time or creates rework nobody planned for. This article describes what text recognition really does, where its limits lie, how to measure recognition quality on your own documents, which fields can be extracted reliably and what setup effort to expect per document type.

Key takeaways

  • Text recognition turns pixels into characters; it does not understand the document. Whether a number is the gross amount or an order reference is decided by field mapping — and that mapping needs its own setup for each document type.
  • The source material determines the result more than the choice of software: digitally created PDF invoices already contain the text, whereas carbon copies, thermal receipts, stamps across printed text and handwriting cause systematic gaps.
  • Recognition quality should be measured rather than believed: a sample of around 200 of your own documents (project experience), evaluated per field instead of per document, shows within hours which fields may be accepted without a query.
  • For incoming invoices four fields carry the benefit: invoice number, invoice date, gross amount and the tax details. Every additional field costs extra setup and returns noticeably less time saved in proportion.
  • A review step by a person remains part of the workflow. It does not disappear with better recognition; it shrinks from retyping to a short comparison of the extracted fields against the document image.

What text recognition actually does

Text recognition is a translation step from pixels to characters. A scanner or a camera delivers an area of light and dark dots; the software looks for lines within it, separates characters from one another and assigns each character the most probable letter or digit. Every one of those decisions carries a confidence value. That value is the most important part of the result, yet in practice it is the part most often passed over.

Meaning does not arise at this stage. A recognised string of characters only becomes an invoice number, a gross amount or a service date through field mapping. That mapping works with positions on the page, with keywords next to a value and with format rules, for example the expectation that an amount has two decimal places. This is why setup is tied to a document type: an invoice from a recurring supplier looks nothing like a delivery note from goods receipt.

The third step is the business check, and it has nothing to do with image recognition any more. Does the amount match the purchase order, does the supplier exist in the master data, has the invoice number already been posted? These checks run against your own systems and can be automated independently of recognition quality. Separating the three steps keeps a project focused on the right questions — otherwise poor source material gets mistaken for poor software.

Two very different files sharing one extension

A digitally created PDF already contains the text as characters; no recognition is needed at all and the values can be taken directly. A scanned PDF, by contrast, contains only an image of the page. Both carry the same file extension and look identical on screen. Before any setup it is therefore worth counting how much of the incoming pile already contains text — that share is the simplest and most dependable part of the project.

The source material decides, not the software

Conversations about document capture usually revolve around which program should be purchased. The more useful question is what the documents arriving every day actually look like. A digitally created invoice and a creased carbon copy are worlds apart, and no software levels that difference. The differences are predictable, which is precisely why they can be examined before the decision rather than after the rollout.

A plain count of incoming documents over two to four weeks helps: how many arrive as digital files, how many as scans, how many as photos from a vehicle or a building site, how many on thermal paper. That distribution says more about the expected benefit than any demonstration. It is also the basis for estimating effort per document type and for setting the order of work.

Source typeTypical starting pointWhat realistically works automatically
Digitally created PDFText present as characters, fixed structure per senderAll relevant fields without a recognition step, very low error rate
Clean scan at 300 dots per inchFed straight, good contrast, no creasesHeader and footer fields dependable, line items with limits
Photo from a mobile phoneSkewed, uneven lighting, partial shadowsAmount and date usually, number fields depending on sharpness
Thermal receiptLow contrast, narrow, fades over timeAmount and date with checks, individual items rarely sound
Carbon copyWeak print, colour cast, background patternOnly individual fields, high share of queries
Document with stamp and handwritingStamp covers print, notes in the marginPrinted fields away from the stamp, handwriting not at all

The order of work follows almost by itself. Setting up a document type with high volume and good source quality first produces a visible effect within days. Starting with the most difficult pile uses up everyone's patience on a special case. This prioritisation is part of a process analysis and takes less time than most people expect.

Where recognition regularly fails

The failure patterns repeat across industries. They are not caused by bad software but by source material that physically no longer carries the information. A character present on paper only as a hint cannot be reconstructed — at best it can be guessed, and guessing is the most expensive mode of operation where amounts are concerned.

The distinction between two kinds of error matters. An empty field is inconvenient but harmless: it is noticed and filled in. A wrongly filled field with a high confidence value is the expensive case, because it travels unnoticed into the posting. A setup that prefers to enter nothing when in doubt makes for calmer daily operation.

  • Carbon copies. The print is weak, the paper often coloured, with a pattern in the background. Digits of similar shape merge, an eight becomes a six or a zero.
  • Stamps across the text. A received stamp or a payment note across the amount line makes the affected characters unusable. The remedy is not technical but an instruction to stamp into free areas.
  • Handwriting. Notes, initials, account codes and signatures are practically unusable for automatic evaluation. Where handwritten details are genuinely needed, they belong in a form or a mobile capture at source.
  • Tables without rules. Line items are separated by spacing rather than borders. If a row slips, the quantity ends up in the unit price. This is why header and footer fields are far simpler than line item data.
  • Creases, staples and folds. They create shadow edges that get read as characters. A document folded twice in an envelope yields visibly worse results than the same document lying flat.
  • Foreign-language documents and special characters. Different date conventions, different decimal separators, differing labels for tax details. Without a rule per language, wrong values appear quietly.

Every one of these points has an organisational lever that is cheaper than any technical fix: file flat instead of folding, stamp into free areas, replace carbon copy sets with a digital form, ask suppliers to send digital files. These measures cost no budget and improve the starting position permanently.

Measure recognition quality instead of trusting it

Vendor figures on recognition rates refer to test sets that have little to do with your own incoming post. The statement only becomes sound with your own documents. A sample of around 200 items (project experience) drawn from live operation across senders and source types is enough for a decision. What matters is that the sample reflects everyday reality rather than the finest specimens.

Measurement is done per field, not per document. A statement such as "96 per cent recognition" is of no help if the missing remainder happens to be the invoice amount. For each field three values are recorded: correctly taken over, left empty, wrongly filled. The third is decisive, because it determines how high the threshold for automatic acceptance has to be. Recognition without your own measurement is a guess with a technical surface.

Terminal
$ documentcheck --sample incoming-invoices --size 200
Sample: 200 documents, 118 digitally created, 82 scanned Field correct empty wrong Invoice number 189 9 2 Invoice date 194 5 1 Gross amount 191 8 1 Tax amount 176 21 3 Supplier match 183 14 3 Order reference 121 71 8 Note: all 15 wrong values came from scanned sources Recommendation: do not auto-accept the order reference

The figures in the example come from a measurement on a mixed pile of documents (project experience) and are not a value to expect for your own incoming post — their purpose is to show the shape of the evaluation. A table like this leads to a clear rule: fields with very few wrong values are accepted automatically, fields with a noticeable share of errors go into the review step without a pre-filled value. The measurement should be repeated after a few months, because senders, forms and paper quality change.

Which fields can be extracted reliably

For incoming invoices a small number of fields carry the bulk of the benefit. They sit in predictable places, follow a checkable format and can be reconciled against your own data. Anything beyond that is possible, but it costs additional setup and returns proportionally less time saved.

These fields have a second advantage: they can be sanity-checked immediately. A date in the future, an amount without decimal places, an invoice number that has already been posted — such checks are simple rules against your own data and catch a large share of recognition errors before a person even looks at the document.

Invoice number

Usually in the header area, often with a clear label. Checkable against numbers already posted for the same supplier, so duplicate postings surface early.

Invoice and service date

Clear format, easy to check against plausibility limits. Take care with foreign-language documents where day and month appear in a different order.

Gross and net amount

Two decimal places, fixed position in the footer area. The cross-check of net plus tax equalling gross is an effective self-control.

Tax amount and tax rate

Frequently spread over several lines with mixed rates. Recognised individually and then checked against the total, otherwise quiet differences appear.

Tax number and VAT identification number

Fixed character patterns and therefore easy to check. It also serves to match the sender to a master data record unambiguously.

Bank details and payment terms

Easy to recognise but security-relevant: a changed bank account does not belong in automatic acceptance but in a deliberate release.

extraction-result.txt
document:         2026-06-24_incoming_00417.pdf
source:           scan, 300 dpi, greyscale

field                value                conf.   acceptance
invoice_number       RE-2026-004182       0.99    automatic
invoice_date         2026-06-24           0.98    automatic
gross_amount         1284.50              0.97    automatic
tax_amount           205.12               0.93    automatic
tax_number           DE...                0.96    automatic
order_reference      (empty)              --      review step
bank_details         DE.. (changed)       0.95    release needed

cross-check:      net 1079.38 + tax 205.12 = 1284.50          ok
duplicate check:  invoice number not yet posted               ok

Line items are the special case. They are appealing because they allow reconciliation with the purchase order, but technically they are the most demanding part. They make sense where many similar documents come from the same sender and the structure is stable. For mixed incoming post they are the wrong starting point. Passing the checked values on to accounting or the merchandise management system is then a matter of data integration and is solved independently of recognition.

Why the human review step stays

The common expectation is that good recognition replaces human control. In fact it shifts what that control consists of. Instead of typing figures, a person compares the pre-filled fields with the document image and releases them. Several minutes of capture become a few seconds of visual check — that is the real gain, not the removal of the step.

This shift needs a screen that shows both side by side: the document image with the source position highlighted on the left, the extracted fields on the right. Anyone who has to switch between two programs to check loses the advantage again. It is equally important that uncertain fields are visibly marked, so attention goes where it is needed.

Four controls that make the review step effective

First, source highlighting: every extracted field points to the place in the image it came from. Second, the threshold: values below the defined confidence are left empty rather than pre-filled. Third, the cross-check: net plus tax has to equal the gross amount, otherwise the document goes to review regardless of confidence. Fourth, the duplicate check against invoice numbers already posted for the same sender. These four controls cost little and catch most of the cases in which a wrong figure would otherwise slip through unnoticed.

The review step also needs a cover arrangement. A workflow that depends on a single person builds up a backlog within days during holidays. Rollout and training should therefore settle who checks, who covers and at what backlog someone is informed.

Text recognition does not replace the check, it replaces the typing. Removing both at once merely moves errors into accounting.

Principle from project work

Effort per document type: what to expect

Setup does not happen once for the company but per document type and in part per sender. For each combination the fields are defined, source positions described, check rules formulated and the threshold for automatic acceptance set. A test run with real documents and a correction round follow. The sequence is manageable, but it repeats for every type.

The overview below gives experience-based figures for setup including measurement and test run (project experience). It does not replace an estimate against your actual pile, but it conveys the order of magnitude and above all shows which share of the work stays with people permanently.

Document typeSetup per typeWhat is still done manually
Invoice from a recurring supplier0.5 to 1 dayVisual check of the pre-filled fields, release
Invoice from changing senders2 to 4 daysCheck with more frequent corrections, maintaining new senders
Delivery note with line items3 to 5 daysLine items on a sample basis, deviations in full
Order confirmation1 to 2 daysReconciliation with your own purchase order
Till receipt on thermal paper1 dayAmount and date are regularly entered afterwards
Handwritten timesheetrecognition not sensibleMove capture to source instead of scanning

On top of that comes the ongoing share: adding new senders, following changed forms, reviewing the thresholds after a few months. This effort is small, but it is not zero. Leaving it out of the operating concept produces recognition that gets quietly worse over a year without anyone being able to name the reason. Such maintenance is part of ongoing IT operations.

Filing, traceability and the legal frame

Recognition does not conclude the case. The document image, the extracted values, any corrections and the note of who released them belong together and must remain findable. Only that connection turns a collection of data into a traceable case: months later it must be visible which value was accepted automatically and which was changed by hand.

The legal frame touches several points at once: retention periods for accounting documents, the unalterability of the filed version, a description of the procedure as well as controlled access rights and a working backup, which counts among the basic requirements for IT systems (German Federal Office for Information Security). In addition, companies in domestic business transactions have had to be able to receive electronic invoices since 1 January 2025 (German Value Added Tax Act) — a structured invoice data set needs no text recognition at all and is therefore the better route wherever it is available.

This description does not replace a legal review

Retention, procedural documentation and the handling of personal details on documents depend on the specific situation of a company. The points here are technical orientation for the conversation with your tax adviser or lawyer, not legal advice. Anyone scanning documents and destroying the paper version afterwards should describe their own approach in writing beforehand and have it reviewed professionally.

An often underrated side benefit of orderly filing is search. As soon as invoice number, sender, date and amount exist as data, a document is findable in seconds rather than in a binder. In many companies this effect is felt more strongly day to day than the capture time saved, because it also shortens queries from customers, suppliers and auditors. How paper and digital filing can be brought together is described on the document digitisation page.

Rollout in four steps

The order matters more than the pace. Starting with measurement rather than with choosing a program means the decision rests on your own documents instead of a demonstration with someone else's material.

Record over a typical period which document types arrive in what quantity and in what form: digital file, scan, photo, thermal receipt, carbon copy. This count runs alongside daily business and takes a few minutes a day.

The parallel run in step 3 is frequently skipped and is nevertheless the cheapest part of the project. It reveals deviations while they are still without consequence, and it removes the worry that a changeover might endanger the monthly close.

When the effort pays off

The calculation is simpler than it is often made out to be. Four figures are needed: the number of documents per month, today's capture time per document, the future check time per document and the one-off setup effort. The difference between capture and check, multiplied by the volume, gives the monthly relief. Without a short timing exercise at the actual desk, every business case remains a wish.

Two items are usually missing from that calculation. First, the error costs of today's manual capture: a transposed figure in an amount, an invoice posted twice, a payment term overlooked. Second, the search effort that disappears with orderly filing. Neither can be quantified to the decimal place, but both are quickly named in conversation with the departments concerned.

  • Measure the time per document before and after, do not estimate it. Two hours of observation at the desk produce sounder figures than an hour of discussion in a meeting room.
  • Only count the document types that will actually change. A pile that stays manual because of handwriting does not belong in the savings.
  • Set the check time honestly. Visual check and release cost time, even when they are short. Putting them at zero produces a calculation that will not hold up in practice.
  • Plan for the ongoing maintenance share, meaning new senders, changed forms and the repeated measurement after a few months.
  • Name the side benefits without pricing them: faster answers to queries, orderly access, less searching through binders.

For orientation on the order of magnitude: a process analysis with counting, measurement and prioritisation starts at 1,900 EUR net, implementation per automation or interface at 4,900 EUR net, ongoing support at 190 EUR net per month. The structure behind that is set out on the pricing page. Whether the step is worthwhile is decided by the measured difference between capture and check, not by the number of features in a piece of software.

This article is based on data from: the German Federal Office for Information Security, the German Value Added Tax Act and our own project experience from digitisation projects in mid-size companies.

Related Articles

KI & Automatisierung

AI in back-office work: the routine it takes off your desk

Concrete, not hype: which back-office routines AI reliably prepares in 2026 - sorting mail, classifying enquiries, reading receipts - and where people decide.

13 min read
Practice & rollout

Replacing the grown spreadsheet: when and how it pays off

When a spreadsheet becomes a shadow core system: five weak points, three decision questions and a transition path with a parallel run instead of a standstill.

14 min read
Data & documents

Document management: from folders to a searchable archive

From a folder of PDF files to a searchable archive: indexing, linking to the case, versions, access rights, retention and how to handle the existing backlog.

14 min read