An update is rarely the problem. The problem is the moment it goes in: in the middle of order intake, days before the month-end close, or on the afternoon when half the workforce is already waiting on a system that has stopped responding. So it gets postponed. Postponement turns into a quarterly rhythm, the quarterly rhythm turns into a backlog, and at some point the question is no longer when the window sits but why a known hole has been open for months. This article describes the opposite approach: a small, fixed slot in the weekly grid, a test instance that runs ahead of it, and a fallback path rehearsed before it is needed. Plus the technology that keeps the restart out of the window, and the records that remain at the end.
Key takeaways
- The pace is set outside your company. During the reporting period of the German security situation report, an average of 119 new vulnerabilities became known worldwide every day (BSI) - one window per quarter cannot catch up with that.
- A fallback path is not optional. The German IT baseline protection catalogue states in basic requirement OPS.1.1.3.A1 that fallback solutions must be available when patches are installed (BSI).
- Separate security fixes, functional releases and platform changes. Only the first class belongs in the short cycle; the other two need testing, notice and a slot of their own.
- Technology moves the restart, it does not remove it. Kernel livepatching is, by the vendor's own statement, not a replacement for rebooting (Canonical) - the restart simply moves into a planned slot.
- Evidence is part of the window. Anyone who has to issue an early warning within 24 hours of becoming aware of a significant incident (Official Journal of the EU) needs the change record beforehand, not afterwards.
The pace is set outside, not by your quarterly plan
Between 1 July 2024 and 30 June 2025, an average of 119 new vulnerabilities in IT systems became known worldwide every day (BSI). That is not an outlier but a trend: compared with the previous reporting period it amounts to growth of around 24 percent (BSI), which the agency attributes only in part to a changed reporting practice. A company that opens a window four times a year is therefore working against a volume that keeps growing between two slots. The backlog is not individual negligence; it follows from the chosen cadence.
How quickly that backlog becomes dangerous is visible in the European threat assessment. In 21.3 percent of the cases evaluated, an exploited vulnerability was the way into the network, and broad campaigns weaponise such flaws within days of disclosure rather than within months (ENISA). The assessment draws on 4,875 incidents from the same reporting period. At the same time, the number of vulnerabilities for which exploitation is actually documented stays manageable: 245 entries were added to the relevant catalogue during the period (ENISA). That list is the work queue to clear before any other.
The economic frame has been surveyed as well. The industry association Bitkom puts the damage caused by cyberattacks in the German economy at 160.4 to 205.8 billion euros (Bitkom); the survey covered 1,003 companies with at least ten employees and at least one million euros in annual revenue. 96 percent of companies were recently affected by data theft, industrial espionage or sabotage, or suspect they were (Bitkom). At the same time only 43 percent still consider themselves very well prepared for cyberattacks, down from 50 percent the year before (Bitkom). The gap between exposure and preparation is exactly where an orderly maintenance window does its work.
What the figures mean for planning
What the rulebooks expect from a patch process
The German IT baseline protection catalogue treats patch and change management as a module of its own. Requirement OPS.1.1.3.A15 states plainly that patches should generally be installed promptly after publication (BSI). A window four months away clearly does not meet that; a window four weeks away does for most cases. The module says nothing about the hour of the day or the duration - it asks for a governed procedure, not a particular night shift. Reading that before looking for the first slot saves the debate about the perfect window.
Two further sentences from the same module matter more for planning than any tooling question. When patches are installed and changes are carried out, fallback solutions must be available (BSI) - that is a basic requirement, not a suggestion. And patches and changes should be suitably tested in advance (BSI). Test instance and fallback path are therefore not a later expansion stage but part of a procedure a company owes anyway. How to record that procedure without writing a manual is covered in our article on process documentation for audits.
At European level the NIS2 Directive adds to this. Article 21(2)(e) explicitly names security in network and information systems acquisition, development and maintenance, including vulnerability handling and disclosure (Official Journal of the EU) - one of ten minimum measures. Installing updates therefore becomes a legal duty for the entities in scope and, through supply chains, a contractual topic for their suppliers. Who is affected in a mid-size company is set out in our article on NIS2 for mid-size companies.
| Rule | What it requires | What it means in the maintenance plan |
|---|---|---|
| IT baseline protection OPS.1.1.3.A15 | Install patches promptly after publication (BSI) | Short cycle for security fixes, separate from functional releases |
| IT baseline protection OPS.1.1.3.A1 | Fallback solutions available, changes tested in advance (BSI) | Test instance ahead of the window, fallback path planned and rehearsed |
| NIS2, Article 21(2) | Maintenance and vulnerability handling as a minimum measure (Official Journal of the EU) | Named role, fixed slots, records anyone can follow |
| NIS2, Article 23(4) | Early warning within 24 hours, report within 72 hours (Official Journal of the EU) | Keep change records retrievable without searching |
| Cyber Resilience Act, Article 13 | Support period of at least five years (Official Journal of the EU) | A floor for procurement, not a target value |
At first glance the reporting deadlines in Article 23 look like an emergency topic. In fact they decide how records are kept: anyone who has to submit an early warning within 24 hours of becoming aware of a significant incident (Official Journal of the EU) and follow up within 72 hours with a report containing an initial assessment (Official Journal of the EU) has no time to reconstruct the patch level. A maintenance plan that records which release landed on which system and when is therefore not paperwork but the basis for being able to answer at all in day-to-day IT operations.
Three kinds of change, three cadences
The most common planning mistake is to put everything into the same window. A security fix for a library, a new functional release of the ERP system and the move to a new operating system version carry different risks, different depths of testing and different fallback routes. Bundling them produces a long window with many people involved - and, when something breaks, the question of which of the bundled changes caused the outage. Separating them costs planning time and saves diagnosis time.
Security fix
Small scope, known trigger, short test. It belongs in the shortest cycle and needs no notice to the business units as long as it works without a restart. The fallback is usually withdrawing a single package.
Functional release
New screens, changed fields, altered reports. Here the test on the test instance decides, and the business units need notice with a date. The fallback also touches data created under the new release.
Platform change
New operating system version, new database release, new runtime. That is a project with its own slot, its own rehearsal and its own fallback path - and the point at which replacing the legacy system deserves a look.
Assigning a change to a class becomes easier once the rule is written down. Three lines are enough: what counts as a security fix, who decides in case of doubt, and which cadence belongs to which class. After that, whether a package may go into Thursday's window is no longer a discussion but a look at the rule. Where the assignment stays disputed, a process analysis that records the path from advisory to installed release pays for itself.
Technology that postpones the restart
For server operating systems, procedures that make security fixes effective without a restart have been available for some years. For Windows Server the vendor describes a quarterly rhythm: during the two months following a baseline release, devices receive a hotpatch update that contains only security updates and can be installed without a restart (Microsoft). A planned year therefore consists of four baseline releases with a restart and eight hotpatch releases without one (Microsoft): eight of the twelve releases can be installed without a restart. A year with only four restart windows does not follow from that. The same vendor states in the same document that non-security updates for Windows, .NET updates and driver and firmware updates sit outside the Hotpatch program and require the machine to be updated during hotpatch months as well; a new baseline adds a periodic restart on top (Microsoft).
The price appears in the same documentation: hotpatch updates do not support automatic rollback (Microsoft). That shifts the burden, it does not remove it. Anyone taking the restart out of the window has to build the fallback path elsewhere - through an image of the system state before the change, through a second instance, or through the ability to withdraw the affected package on its own. This decision belongs in the plan before the first hotpatch runs.
On Linux, kernel livepatching plays the same role. The vendor of the widely used distribution states coverage of ten years for its service, with an add-on extending this to fifteen years (Canonical). The same page is equally clear that livepatching is not a replacement for rebooting (Canonical). That is precisely the operational gain: the restart does not disappear, it moves out of time pressure into a planned slot when somebody is on site anyway.
Where services run in several instances, the question shifts again. In a rolling update of a workload, at most 25 percent of the instances are unavailable at the same time by default (Kubernetes); the rest keep serving requests. That is not an argument for a new platform but a pointer to the pattern behind it: several identical instances, a dispatcher in front of them, and a check that returns an instance to service only once it responds. The same pattern works without a container platform - two application servers behind a dispatcher are enough.
Where the technology stops
The weekly grid: a window operations can carry
A maintenance window usually goes unused not because it is too small but because it is too vague. A fixed weekday at a fixed time beats a monthly slot renegotiated every time: the business units know when not to expect a report, the on-call staff know when to watch, and the change that did not get finished waits for next week instead of next quarter.
- Monday to Wednesday: the new release runs on the test instance. Testing happens on real cases, not on an empty database.
- Wednesday: notice to the business units with date, duration and the functions that may be affected. One line is enough, it only has to arrive.
- Thursday before the window: take the backup and evidence a sample restore, rather than only looking at the backup log.
- Thursday inside the window: install in a defined order, one section at a time, with a check after each section.
- Thursday after the window: a short functional test on a handful of real workflows, documented with time and result.
- Friday: follow-up. Collect feedback, record open points in the case file and plan the next cycle.
The order inside the window matters more than its length. Touching the database first, then the application and the interfaces last leaves a check after every step that narrows down the fault. Starting everything at once saves twenty minutes and loses them twice over when something breaks. The written plan captures exactly that order - with an abort point per section at which the fallback begins instead of the search continuing.
Window: Thursday 22:30 to 23:15, calendar week 38
Release: business application 12.4.1 -> 12.5.0
Mon 09:00 release installed on the test instance
Tue 14:00 test run on real cases, result filed with the case
Wed 10:00 notice to business units (duration, affected reports)
Thu 21:30 backup taken, sample restore evidenced
Thu 22:30 section 1 database -> check 1, abort point A
Thu 22:45 section 2 application -> check 2, abort point B
Thu 23:00 section 3 interfaces -> check 3, abort point C
Thu 23:15 functional test on real workflows, sign-off or fallback
Fri 08:00 follow-up, feedback, open points
Fallback: abort points A to C, each with order and ownership
Records: case number, releases, times, check resultsA plan like this fits on one page and is merely updated the second time round. Its real value is that it answers, in advance, the questions nobody wants to answer inside the window: who decides on the abort, who is reachable, which credentials are needed and where the backup sits. That belongs in the same place as the rest of the process documentation, not in an inbox.
The fallback path is the part that must be rehearsed
That the fallback is the weakest part shows even where supervision and audits have long been in place. Among operators of critical infrastructure, around 80 percent already ran an information security management system with a maturity level of at least three, while the share for business continuity management systems was considerably lower at just under two thirds (BSI). The intent is there; the rehearsed ability to fall back lags behind. In smaller companies the gap tends to be wider, not narrower.
A fallback path consists of four statements: the state you return to, the route there, the person who triggers it, and the moment the decision is taken. Without the last one, people keep searching inside the window until the window is over. That is why the plan carries an abort point per section. And that is why the route is walked through once a year - in a slot without time pressure, using the same backup that would be used in earnest.
A backup that has not been restored is an assumption
- Target state named: version, database release, configuration release
- Route described: order of the steps, credentials required, estimated duration
- Decision governed: who aborts, from what point in time, with which deputy
- Backup verified: restore test with a date, not merely a green log
- Data considered: what happens to cases created after the change
- Communication prepared: who is informed, by which route, with which wording
- Rehearsal scheduled: one dry run a year, documented like a real window
The test instance: small, but built the same
A test instance does not need the performance of the production system, but it does need the same structure. A different runtime version, a different database release or missing interfaces turn the test into reassurance without evidential value. Where the data volume rules out copying, an extract with real structures helps: one tenant, one period, a set of typical cases. Personal data is replaced before it leaves the production environment - a point our article on security for interfaces picks up as well.
The second point is who uses it. A test instance seen only by IT tests technology. A test instance where two people from the business unit work through real cases for an hour tests the workflow. The difference shows up in small things: a field that changes position in the new release, a printout that loses a line, a report that rounds differently. Those findings no longer surface inside the window; they surface the following Monday - at the customer.
The most expensive maintenance window is the one in which the fallback gets tried for the first time. Afterwards nobody mentions the two hours a rehearsal would have cost.
Lifecycle: what procurement settles in advance
Part of the maintenance burden is created years before the first window - during selection. The European Cyber Resilience Act obliges manufacturers of products with digital elements to a support period of at least five years (Official Journal of the EU). Security updates that have been provided must also remain available for at least ten years after provision (Official Journal of the EU). For procurement that is a floor, not a target: anyone planning to run a system for ten years asks for the period contractually.
A second requirement in the same regulation looks unremarkable and changes maintenance planning considerably. Where technically feasible, new security updates must be provided separately from functionality updates (Official Journal of the EU). That separation is the precondition for running a security fix in a short window without the business units finding new screens afterwards. Anyone writing a tender today can include the separation as a requirement instead of missing it later.
The deadlines of the regulation are already fixed. Article 14 applies from 11 September 2026 (Official Journal of the EU), the regulation otherwise from 11 December 2027 (Official Journal of the EU). For actively exploited vulnerabilities a final report is due at the latest 14 days after a corrective or mitigating measure becomes available (Official Journal of the EU). What that means for your own reporting route is described in our article on the 24-hour reporting process.
| Component | Support according to the publisher | What follows for the plan |
|---|---|---|
| PHP branch, full support | two years from the initial stable release (PHP Group) | Plan a version change every two years rather than on demand |
| PHP branch, entire period | four years, after which support ends (PHP Group) | The latest date for the change is fixed on the day of rollout |
| Windows 10, Home and Pro | end of support on 14 October 2025 (Microsoft) | Workstations without a successor release belong in a project of their own |
| Products with digital elements | support period of at least five years (Official Journal of the EU) | A floor for the tender, contract length to be settled separately |
| Security updates | available for at least ten years (Official Journal of the EU) | Catching up older systems stays possible; plan where to keep the sources |
The table is not a procurement plan, it is a calendar. Every row carries a date that is already fixed today, and every one of those dates creates work that can either be spread out or not. A company that collects the end dates of its components once a year turns later emergencies into ordinary projects. Where systems appear for which no successor release exists, the work on replacing the legacy system starts earlier than planned, but not as a surprise.
Records that justify the effort
A maintenance window produces documents, and those documents are the part noticed outside IT. They answer three questions that get asked sooner or later: which release ran when, who approved it, and how long a known gap stayed open. Anyone keeping those three items per change has an answer for auditors, insurers and customers - and, internally, a basis for deciding the next cadence in IT operations.
- Case number per change, with date, window and the systems involved
- Source and target release per system, taken from the asset inventory
- Sign-off with name and time, even when it happens in a single line
- Result of the functional test, listing the workflows checked instead of a tick
- Time between publication of the fix and installation, measured per change
- Deviations and aborts, with the reason and the point derived from it
The fifth line is the only key figure this procedure really needs. It says how long a known hole stays open in your own house, and it can be kept without any tooling. If it falls from weeks to days over a year, the window has done its job. Where change and fault notices currently pile up in one shared inbox, our article on the way from a shared mailbox to a traceable case is worth a look; and if machine data is to feed the reporting, the article on access to machine data sets out the new entitlements.
Sources and studies
Related Articles
Own server or data centre: making the decision soberly
Cost over five years, availability, responsibility during incidents, data protection and getting data back — how mid-sized firms decide where their servers run.
Verification of payee in payment runs: handling mismatches
Since October 2025, banks check name and IBAN before every credit transfer. How supplier master data, payment blocks and call-backs handle the bank's responses.
Interim Payments and Variations on Building Sites
How a stage of completion becomes an interim invoice, why a variation is a case with a deadline of its own, and how that produces a final account that can be checked without rework.