Backups are one of those topics where almost every company nods with a clear conscience. Software is running, there is a schedule, and every morning a green tick appears in the report. The uncomfortable questions only arise in an emergency: how long does it take until order processing runs again, and how much work is lost by then? Anyone unable to answer both questions has backup software rather than a backup strategy. This article describes what a sound backup looks like in a mid-size company: several copies, at least one off site, at least one without a permanent connection to the network, and regular restore tests with measured durations. It also covers what encryption attacks change about that calculation, because they now target the backup first.
Key takeaways
- A backup only becomes a backup through a proven restore: as long as nobody has restored a data set in a separate environment and checked it from a business point of view, its usability is an assumption rather than a finding.
- Two figures govern everything else, and both come from the business: the tolerable downtime per workflow and the tolerable data loss, meaning the period whose work would have to be done a second time.
- The minimum is several copies on different systems, at least one of them at another location and at least one without a permanent network connection (BSI), because fire, theft and encryption otherwise hit the original and the copy together.
- Encryption attacks shift the focus from the backup run to retention depth: attackers often move through a network unnoticed for weeks and delete reachable backup states before encrypting, which is why older and immutable states remain necessary.
- Recovery consists of more than data: it needs a sequence, credentials, replacement hardware and named responsibilities, and those details belong in a short plan that stays readable even when no system in the building is available.
Why the backup itself is rarely the actual problem
In most companies where we take stock, the backup is running. It often runs reliably and has done so for years, and the daily report says success. What that report confirms, however, is only one thing: data was written. It says nothing about whether the written data is usable, whether all required systems are included, whether the data set can be retrieved without the very software that has just been encrypted, and how long restoring it would take. Those four points decide the outcome in an emergency, and none of them is covered by the green tick.
The findings resemble each other more than one would expect. Databases of the core application are frequently backed up while running, so the files exist but are not internally consistent. The backup target frequently sits on the same network, under the same credentials and on the same power circuit as the source. Mailboxes are frequently missing because they are hosted by a provider and therefore count as somebody else's problem. And there is almost always no record of who starts what, in which order, in an emergency. Day-to-day IT operations keep the systems alive, but preparing for their failure is a separate task.
On top of that, the emergency is rarely spectacular. A project folder deleted by accident, a failed disk, an update of the core application that goes wrong, a device that no longer starts after a power cut: these are the everyday cases. An encryption attack is the exceptional case with the largest impact, but the preparation for the everyday cases is the same. A backup that has never been restored is not a backup but an assumption.
A green tick confirms a write operation
Two figures that come from the business
Before anyone talks about media, schedules and retention, two decisions are needed per workflow. The first is the recovery time: how long may this workflow stand still before the damage exceeds what the company can bear? The second is the tolerable data loss: how much work may be lost and have to be recorded again? Both figures come from the business and not from the technology. Technology derives the backup schedule, the medium and the procedure from them, but it does not set the values itself.
Setting them works best in conversation with the departments and along individual workflows rather than along servers. The question is not how important a system is, because then everything is important. The question is: what happens in concrete terms during the first four hours without this system, what happens on the second day, what happens in the second week? The answers produce a ranking, and that ranking is the real yield of the exercise. Anyone taking stock for a process analysis anyway can settle both questions in the same meeting.
| Workflow | Downtime bearable up to | Data loss bearable up to | What follows from it |
|---|---|---|---|
| Order intake and inventory management | a few hours | the last completed morning | backup several times a day, restore path prepared and rehearsed |
| Accounting and payments | one working day | one working day | daily backup, longer retention because of record-keeping duties |
| File storage and project folders | one working day | a few hours | versioned states so single files return without a full restore |
| Mailboxes and calendars | a few hours | a few hours | a separate backup, even if the mailbox is hosted by a provider |
| Payroll | until the next payroll run | the last completed run | backup before and after each run, tightly restricted access |
| Technical drawings and machine data | one to two working days | one working day | plan for large volumes, measure the restore duration |
The fourth column shows why this exercise is not academic. A tolerable data loss of a few hours necessarily implies a backup interval shorter than that period. A recovery time of a few hours implies that the fastest copy has to be on site, because retrieving large volumes over a connection takes longer than the business can wait. Together, the two produce a backup architecture that can be justified instead of being carried forward from the past.
Two figures that determine everything else
Several copies, one off site, one disconnected
A simple rule has become established as the minimum: several copies of the data on at least two different systems or media, at least one of them at another location (BSI). The rule looks banal, but it answers exactly the three kinds of failure that actually occur in mid-size companies. The second copy protects against a defective medium or device. The other location protects against events that hit the entire room: fire, extinguishing water, burst pipes, burglary, lightning. And the third addition, which has become just as important, protects against an attacker inside the network: at least one copy that is not permanently reachable.
This third requirement is the one most often overlooked, because it is inconvenient. A permanently mounted network share, a constantly attached drive, a backup target using the same credentials as the production system: all of that is convenient, and all of it is encrypted along with the source during an attack. A copy that is permanently reachable shares the fate of the system it is meant to protect. The distance has to be enforced technically rather than maintained by carefulness.
- Separate credentials for the backup. The backup system does not belong in the company's central user directory. Whoever takes over the administrative accounts of daily business takes over the backup with them.
- Pull instead of push. If the backup system actively fetches data from the source systems, no production system needs write access to the backup target. The opposite direction turns every infected source system into a route to the archive.
- Use retention locks. Storage with immutable retention does not allow deletion or overwriting before the period expires, not even by an administrative account. It is the single most effective building block against deleted backup states.
- Rotate removable media. A set of drives connected in turn and stored off site is unfashionable, inexpensive and unreachable during an attack. For small companies it remains a sound solution.
- Stagger retention periods. Daily states for everyday use, weekly and monthly states in case damage is noticed late. Short retention no longer catches a data error that went unnoticed for weeks.
There are several workable routes for the off-site copy: a second company location, a safe deposit box for removable drives, the data centre of a service provider. What matters is less the route than checking two points. First: is the copy encrypted, and is the key kept separately from it in a place that stays reachable in an emergency? Second: how long does it take to retrieve the data from there? With larger volumes, shipping a drive is sometimes faster than the line, and that is worth knowing beforehand rather than discovering during an incident.
What belongs in the backup scope and regularly goes missing
The scope of a backup rarely grows along with the company. It was defined once, usually when a server was set up, and since then new storage areas, applications and services have appeared that nobody added to the list. A complete list of what is backed up, with interval, retention and target for each item, is therefore the second foundation alongside the figures from the business. It is also the point where gaps can be found most cheaply: at a table rather than during an emergency.
Core applications and databases
Databases need a backup that produces an internally consistent state instead of copying files while the system runs. This also covers clients, templates and report definitions, which often sit in separate directories.
File storage and project folders
Network drives, folders on individual machines, storage areas set up by departments that nobody registered. That last group is regularly rediscovered during a stocktake and often holds current work.
Mailboxes and calendars
Mailboxes hosted by a provider remain the company's responsibility. Their recycle bin is not a backup, and the provider's recovery windows are usually shorter than a record-keeping duty requires.
Configurations and network structure
Settings of the firewall, network switches, telephone system, time recording and machine controls. Without them the data set is ready after an incident, but nobody knows what the surrounding environment looked like.
Access, licences and keys
Licence records, certificates, credentials for portals and the keys of the backup itself. These details belong in a place that stays reachable even when the company's own systems are down.
Externally hosted services
Customer and supplier portals, storage at service providers, reporting tools. Check for each service which recovery is contractually promised and within which period a deleted data set can be retrieved.
Older applications running on an ageing operating system whose vendor can no longer be reached deserve particular attention. Here the data alone is not enough, because the program needed to read it is missing. In such cases a complete image of the system belongs in the backup scope, together with a note on how it can be started. If a legacy system migration is due anyway, the backup scope is a good occasion to make the dependencies visible.
# Extract from a scope list, one line per item
#
# System / storage Interval Retention Target Tested on
Inventory system (DB) hourly 30 days NAS + external archive 12.01.2026
Inventory system (files) daily 90 days NAS + external archive 12.01.2026
Accounting (DB) daily 10 years External archive 05.02.2026
File storage /projects hourly 90 days NAS + removable drive 09.04.2026
Mailboxes daily 365 days External archive 09.04.2026
Time recording (config) weekly 90 days NAS pending
Firewall (config) on change 5 states External archive pending
CAD storage /engineering daily 180 days NAS + removable drive pending
#
# Column Tested on = date of the last successful restore,
# not the date of the last backup run.The restore test: procedure, interval and log
A restore test is more than opening a backup file. It means restoring a data set in a separate environment, having the associated application open it, checking the content from a business point of view and measuring the time required. The business check is the part most often missing: it is not the IT staff but the person from the department who can tell whether the last order from the previous day is present, whether prices are correct and whether attachments still hang on the right record.
The interval depends on the importance of the system and on how often it changes. For central systems a test at least twice a year has proven workable, for everything else at least once a year (project experience). In addition, a test belongs after every significant change: after an update of the core application, after switching the backup target, after moving a server and after adding a new system to the scope. Without that trigger, a broken backup chain is only noticed at the next regular test.
The second run shows the real benefit of the test. The fault was not in the backup but in the scope list, and without the test it would have stayed hidden until an incident. That is exactly why the log matters more than the result: date, system tested, state restored, measured duration, who carried out the business check, which deviation occurred and which action follows from it. These logs are at the same time the evidence a customer, an insurer or a supply chain review will want to see.
The test does not measure whether the backup worked. It measures how long the business stands still in an emergency, and compares that time with what the business can absorb.
What encryption attacks change about the calculation
Germany's federal agency for information security (BSI) has for years named attacks with encryption malware as one of the most significant threats to companies, explicitly including small and mid-size ones (BSI). Bitkom puts the annual total damage to the German economy from theft, espionage and sabotage in the hundreds of billions of euros (Bitkom), and the digitisation surveys of the DIHK regularly show IT security as a growing challenge for mid-size companies (DIHK). For backups this implies no new technology, but a different weighting.
The first difference concerns the timeline. An attack does not begin with the encryption but weeks earlier with access that is initially only observed and extended. During that time backups are located, rights are taken over and reachable states are deleted or encrypted along the way. Anyone retaining only a few days may be left with states that are already compromised. Staggered retention and at least one immutable or disconnected state are therefore not a refinement but the precondition for being able to go back to a clean point at all.
The second difference concerns recovery itself. After an attack, data is not restored into the existing environment, because its condition is unclear. As a rule the environment is rebuilt and the data set is then loaded after checking. That extends the recovery time considerably compared with a simple disk failure and has to be reflected in the plan. Also to be planned for: the incident has to be assessed, and reporting duties as well as involving the authorities have to be examined. Assessing the individual case in legal terms belongs in qualified hands.
Checkpoints against an attack on the backup
Recovery is more than restoring data
When every screen is dark in an emergency, it is not the quality of the backup alone that decides, but whether somebody knows what has to be done in which order. A recovery plan is not a thick manual. A few pages describing the sequence are enough, provided they are available outside the company's own systems. Paper in a folder, a printout at management level, a copy at a second location: a plan that only sits on the encrypted network drive is not available at the decisive moment.
- Who decides and who informs. A named person with a deputy, plus the rule on who informs customers, suppliers and staff, and from what point onwards.
- Contact details outside the systems. Phone numbers of the service provider, the vendor of the core application, the telephone system and the insurer, written on paper rather than kept in a mailbox.
- Order of recovery. Network before servers, directory service before core application, core application before reporting. The order follows from dependencies, not from the urgency claimed by individual departments.
- Credentials and keys. The credentials for the backup and the encryption keys, kept in a place that remains reachable without the company's own systems.
- Interim means and fallback operation. Which hardware can be procured, within what time, and how the company keeps working in the meantime: paper forms, printed lists, limited acceptance of new orders.
- Return to normal operation. When operations count as restored, which rework is outstanding and who records the cases that arose during the outage.
This plan is not a special task for IT but part of the process documentation. It is written once, briefly reviewed at every restore test and updated whenever the system landscape changes. For it to be useful, it is enough that a person familiar with the site understands it without needing access to a system that has failed.
Retention, records and data protection
A backup is not an archive. It serves to restore an operating state and is overwritten once its retention expires. An archive serves to keep individual documents unchanged and findable over years. In practice the two are often mixed, with two unpleasant consequences: documents subject to retention duties exist only in backup states that nobody can search in a targeted way, and at the same time backups keep growing because nobody dares delete older states. Separating the two tasks saves storage and creates clarity.
Documents subject to commercial and tax retention duties come with periods and with requirements for traceability and immutability; the process documentation describes how those requirements are met in the company. The backup is one chapter within it, but it does not replace orderly retention of the documents. How far individual duties reach in a specific case depends on legal form, sector and type of document, and should be examined by a qualified adviser. This article provides orientation and does not replace legal advice.
In data protection terms a second question arises: how does a deletion request fit together with backup states that still contain the record? Common practice is to carry out the deletion in the production system, not to edit the backup state individually, to describe the approach in the deletion policy and to make sure that a restored state does not quietly bring deleted data back into operation. Here too, the arrangement and its assessment in the individual case belong in qualified hands, but the procedure should be prepared technically nonetheless.
Four steps to a tested backup
Step 1: Set the figures per workflow (half a day to one day)
Go through the most important workflows with management and the departments and record two values for each: tolerable downtime and tolerable data loss. The result is a short table with a ranking that subsequently justifies every technical decision.
Step 2: Review scope and distance (one to two days)
Draw up a complete list of the systems, storage areas, configurations and access details covered and compare it with the ranking. In doing so, check where copies are held, which of them is permanently reachable and with which credentials it is written. Gaps are sorted by urgency, not by effort.
Step 3: Test and measure the restore (one day per system)
Restore a real state in a separate environment, start the application, have a person from the department check it and measure the total time to a working state. The measured time is compared with the figure from step 1. Deviations are the actual result.
Step 4: Write the plan and set the interval (ongoing)
A recovery plan of a few pages, stored outside the company's own systems, plus a fixed test interval with dates and named responsibilities. Every test produces one line in the log. After changes to the system landscape the scope is updated, otherwise the list is out of date within a year.
Go through the most important workflows with management and the departments and record two values for each: tolerable downtime and tolerable data loss. The result is a short table with a ranking that subsequently justifies every technical decision.
Draw up a complete list of the systems, storage areas, configurations and access details covered and compare it with the ranking. In doing so, check where copies are held, which of them is permanently reachable and with which credentials it is written. Gaps are sorted by urgency, not by effort.
Restore a real state in a separate environment, start the application, have a person from the department check it and measure the total time to a working state. The measured time is compared with the figure from step 1. Deviations are the actual result.
A recovery plan of a few pages, stored outside the company's own systems, plus a fixed test interval with dates and named responsibilities. Every test produces one line in the log. After changes to the system landscape the scope is updated, otherwise the list is out of date within a year.
The effort for these four steps is manageable and largely one-off. What remains is half a day per test run and a short update whenever something changes. Measured against what several days without order processing cost, this is one of the cheapest safeguards a company can put in place. And unlike many other precautions it delivers a verifiable result: a measured time and a log that can be shown to somebody.
The order matters here. Anyone starting with the technology buys storage and software and still does not know afterwards whether the way back succeeds within the time the business can absorb. Anyone starting with the two figures per workflow has a yardstick against which every later decision can be measured, and it usually turns out that part of the existing means is already sufficient and simply has to be used differently.
Related Articles
Cyber Resilience Act: a reporting process in 24 hours
From 11 September 2026 the reporting duty in Article 14 applies. Who reports to whom, what happens in the first 24 hours and which records remain at the end.
Account Access When Staff Join and Leave the Company
How access is ready on the first working day and reliably ends after someone leaves: taking stock, roles, a trigger from the HR system, annual review.
Interface security: accounts, keys and permissions
Sign-in, key handling, encryption in transit, minimal permissions, separate test accounts and key rotation: what to settle for every interface you run.