Every business has work nobody misses: opening and routing incoming mail, sending an enquiry to the right person, preparing receipts for accounting, typing the same paragraph into a reply for the third time. Exactly this routine has taken centre stage since 2026, when according to Bitkom 41 percent (Bitkom) of companies now use artificial intelligence - up from just 17 percent (Bitkom) the year before. The jump affects administration and simple back-office work far more than production. This article stays deliberately concrete rather than promotional: which back-office routines AI reliably prepares today, where people have to decide, and what a rollout looks like that cuts manual work without handing control to a black box. How this fits into an orderly workflow is set out under process automation.
Key takeaways
- AI does not replace back-office work in one move; it takes over individual repetitions: sorting, assigning, reading out, drafting. The case stays with a person, while the dull groundwork moves to the model.
- The benefit is greatest where a lot of similar text work occurs - exactly the area where, according to Bitkom, 71 percent (Bitkom) of AI-using companies already start.
- Well suited is work with a clear pattern and a good data basis; poorly suited is anything involving judgement, liability or a rare exception. Drawing that line is the real project work.
- Control comes from a check step and a confidence value: what the model assigns confidently runs on, what stays uncertain goes to a person. Every suggestion is traceable and correctable (project experience).
- The tool comes after the stocktake: first record the workflow as it is lived, then pick one tightly defined routine as a pilot - not the whole administration at once.
Why AI hits back-office work in particular
The figure shaping every conversation in 2026 is a jump: from 17 percent (Bitkom) to 41 percent (Bitkom) of companies using AI, plus a further 48 percent (Bitkom) planning to. The survey rests on 604 (Bitkom) companies with 20 or more employees, so it looks at mid-size businesses, not corporations. What stands out is not the jump itself but where it lands. Unlike earlier waves, when automation mainly touched machines and warehouses, the current one reaches right into desk work.
The reason lies in the nature of the tools. Language models process text, and back-office work consists largely of text: emails, forms, receipts, notes, replies. Bitkom names text work at 71 percent (Bitkom) as the most common area of use, followed by customer service at 42 percent (Bitkom). Both are the core business of any administration. Where a fixed rule once had to be programmed for a computer to take on a task, a well-chosen example often suffices today - which lowers the threshold for small businesses considerably.
That also explains why many businesses are uneasy at the same time. A system that processes text freely feels more opaque than a fixed rule whose effect is known in advance. This is exactly where this article aims: not whether AI can phrase things impressively, but which manual back-office work it reliably prepares in practice - and where the person keeps the decision. Which workflows are even worth tackling first is settled beforehand by a sober process analysis.
What AI reliably prepares in back-office work
The key word is prepare, not complete. In back-office work AI produces a suggestion, not a conclusion. It sorts, assigns, reads out and drafts - and puts the result in front of a person for approval. Within this division of roles, six routines are dependable today that occur in almost every office.
Sort incoming mail
Incoming emails and scanned post are pre-sorted by topic and ownership and routed to the right place. People only check the uncertain cases instead of every single one.
Classify enquiries
An enquiry is sorted by concern, urgency and the right handler. That replaces the daily manual reviewing and distributing and speeds up the first response.
Read receipts
From an invoice, delivery note or form, amount, date, number and line items are recognised and lifted into fields - the basis for the later check by accounting.
Transfer data
Recognised details are carried into the leading system instead of being retyped. That defuses exactly the duplicate entry that costs the most time and errors day to day.
Draft replies
For recurring concerns a draft reply is produced in the tone of the business. It is read, adjusted and approved - never sent unchecked.
Summarise
Long cases, minutes or email threads are shortened to the essentials so the handler gets an overview faster. The source stays visible alongside.
Two of these routines deserve a closer look because they show effect fastest in practice. When reading receipts, text recognition meets a clear benefit: instead of transferring every figure by hand, the handler only checks the suggested result. What is realistic here and what is not is described in the piece on text recognition in practice. With classification, the gain comes not from a single fast assignment but from the fact that the ever-repeating distribution work is off the desk - time that experience shows is missing in the demanding cases (project experience).
An honest lower bound matters: AI is strong as long as a pattern exists and the data basis holds. With a poorly scanned receipt, a contradictory entry or wording that has never occurred before, reliability drops. That is why every one of these routines needs a path on which uncertain cases visibly go to a person - more on that below. Anyone wanting to end duplicate entry in principle will find the non-AI routes in the piece on eliminating duplicate data entry.
Where people have to decide
As clearly as AI takes on the groundwork, the decision stays with the business. That is not a cautious phrase but a practical limit. A model can make a suggestion, but it carries no responsibility, it is not liable, and it knows the individual case only as far as it appears in the text. Four areas therefore belong explicitly in human hands - not as an exception, but as a fixed part of the workflow.
The first is approval with legal or financial effect: authorising a payment, confirming an invoice as correct, answering a contract. The second is judgement calls where there is no right answer in the data: a goodwill gesture, a discount, handling a complaint. The third is exceptions that deviate from the known pattern - exactly the cases where a model is, in experience, most confidently wrong. The fourth is every case below the set confidence threshold: if the system is not sure enough, it does not guess, it hands over.
Taking the black-box worry seriously
This division of roles has a pleasant side effect: the demanding work gains time, the dull work loses it. Whoever hands off the distributing and retyping can look at the disputed case more closely. That is the real justification for the project - not staff cuts, but a better distribution of the hours already there. Whether the effort pays off can be estimated in advance; the method is in the piece on when automation pays off.
Well suited, poorly suited: a way to sort it
Whether a routine suits AI can be read in advance from a few markers. The comparison below is the short version of what we check first in projects, before a tool even comes into question. It does not replace a conversation, but it sorts reliably.
| Marker | Well suited to AI | Better with a person |
|---|---|---|
| Pattern | Recurring, with clear structure | One-off, different every time |
| Frequency | Many similar cases per week | Few, but heavy cases |
| Data basis | Complete and readable | Gaps, contradictions, poor scans |
| Cost of an error | Small and easy to correct | Legally or financially significant |
| Type of decision | Assign, recognise, pre-draft | Weigh up, goodwill, judgement |
| Checkability | Result quickly reviewable | Assessment needs expertise |
The table also makes clear why all-or-nothing rarely makes sense. Hardly any case falls entirely into one column. As a rule a workflow can be split: the recurring, well-structured share goes to AI, the weighing-up share stays with the person. This split is exactly the work that precedes choosing a tool - and it decides success or disappointment far more than the choice of model. Which processes to tackle first is covered in the piece on which processes to tackle first.
A rollout without losing control
The path from idea to a dependable routine is unspectacular and can be described in five steps. What matters is that the check step is built in from the start and not meant to be retrofitted later.
Step 1: Pick one narrow routine
Not the whole administration, but one well-defined, high-frequency task: pre-sorting incoming mail, say, or reading out one particular type of receipt. Starting small is not hesitation but the precondition for a sound assessment (project experience).
Step 2: Set the confidence threshold
Before go-live it is decided from which confidence value a suggestion runs on without a query and from where it goes to a person. This threshold is a deliberate choice of the business, not a default of the tool.
Step 3: Test on real cases
Testing runs on real cases from recent weeks, not on smooth examples. Only there does it show where the model is sure and where it is not - and whether the threshold sits right (project experience).
Step 4: Go live with support
For the first weeks the routine runs under a double eye: a person sees every suggestion before it takes effect. Only once the hit rate holds are confident cases let through and the checking concentrated on the uncertain ones.
Step 5: Measure and adjust
After about four weeks it is reviewed: how often was the suggestion right, how often did it need correcting, where did handovers pile up? From that picture the threshold is readjusted and the next routine chosen (project experience).
Not the whole administration, but one well-defined, high-frequency task: pre-sorting incoming mail, say, or reading out one particular type of receipt. Starting small is not hesitation but the precondition for a sound assessment (project experience).
Before go-live it is decided from which confidence value a suggestion runs on without a query and from where it goes to a person. This threshold is a deliberate choice of the business, not a default of the tool.
Testing runs on real cases from recent weeks, not on smooth examples. Only there does it show where the model is sure and where it is not - and whether the threshold sits right (project experience).
For the first weeks the routine runs under a double eye: a person sees every suggestion before it takes effect. Only once the hit rate holds are confident cases let through and the checking concentrated on the uncertain ones.
After about four weeks it is reviewed: how often was the suggestion right, how often did it need correcting, where did handovers pile up? From that picture the threshold is readjusted and the next routine chosen (project experience).
The core of this approach is the confidence value. Every result carries a statement of how sure the system is. If it is above the agreed threshold, the suggestion runs on and is spot-checked; if below, the case lands with a person. That keeps the number of manual steps steerable without losing control. A simple log makes it traceable - similar in structure to the extract below.
Inbound | AI suggestion | Confidence | Result
--------+---------------------------+------------+-------------------------------
Mail 1 | Invoice -> accounting | 0.97 | taken, spot-check ok
Mail 2 | Complaint -> sales | 0.91 | taken
Mail 3 | Quote request -> sales | 0.62 | below threshold -> to person
Mail 4 | Application -> HR | 0.95 | taken
Mail 5 | unclear (two topics) | 0.48 | below threshold -> to person
Threshold: 0.85. Of 5 inbound items, 3 ran on confidently, 2 went to a
person for checking. No assignment happened unnoticed - every suggestion
is logged with a confidence value and a source.Data protection and traceability
As soon as AI processes incoming mail, enquiries or receipts, personal data is almost always involved - names, addresses, cases. The General Data Protection Regulation therefore applies, regardless of how new the tool is. In practice that means three things: it must be clear where processing happens, what the data is used for, and who gets to see which suggestion. We rely on processing with hosting and data in Germany and on a log that makes every assignment traceable. The concrete duties in a project are described in the piece on data protection when digitising processes.
A second point concerns employees. A system that assigns inbound items and logs handling can in principle also map behaviour and performance. Where a works council exists it therefore belongs in the process early - not just before go-live, but before the selection. What to keep in mind is set out in the piece on involving the works council in IT projects. Likewise, before go-live there should be a clear statement of what the logs are used for and what expressly not; assessing the legal position in a specific case remains a matter for professional advice, and this article is no substitute for it.
Traceability is not an add-on but the condition
Where a business starts
The larger part of a successful rollout lies not in the tool but in the preparation - and a business can do that itself before a provider is even in the room. Whoever brings one tightly defined routine, a handful of real cases and a sense of the confidence threshold starts under far better conditions. None of this needs outside help.
- Pick a single, frequent routine - sorting incoming mail, say, or reading out one type of receipt - and not the whole administration at once.
- Lay out ten to twenty real cases from recent weeks, deliberately including a few difficult and atypical ones.
- Decide which decision in this routine must never be made without a person - that is the line at which the confidence threshold sits.
- Clarify where the data may be processed and which personal details are involved.
- Name a responsible person who checks suggestions and feeds corrections back - so the system learns from mistakes rather than repeating them.
- Set a date for the review after four weeks before the next routine is added.
From there it can be extended step by step - to the next type of receipt, the next inbound channel, later to automatically generated reports. Anyone framing it more broadly and letting departments build small workflows themselves will find the entry point in the piece on low-code in the business. And anyone coupling back-office work to invoicing should keep the timeline of the e-invoicing mandate 2027 in mind. In all cases the same order applies: workflow first, then tool - and the check step stays.
Related Articles
Handling complaints digitally: deadlines and evidence
Which periods run from delivery, which records should be created on a complaint case, and how to map both digitally without turning it into a large project.
Retaining company knowledge before experience leaves
Capturing head knowledge, documenting critical workflows and testing the stand-in before an experienced colleague leaves: schedule, metrics and sources.
Account Access When Staff Join and Leave the Company
How access is ready on the first working day and reliably ends after someone leaves: taking stock, roles, a trigger from the HR system, annual review.