Most companies do have monitoring in place — it simply answers the wrong question. One service reports that the server is reachable, a second that disk space is sufficient, a third that the backup completed. All three can show green while not a single order has moved from the inventory system to the accounting system for four days. None of those checks looks inside the transaction. Anyone running an interface, or having one run for them, needs monitoring that knows the business rather than only the technology: did any orders arrive today? Do the volumes match what the other side sent? Is anything stuck that was flowing yesterday? This article describes how to build such business-level checks, how to set thresholds without drowning in false alarms, why an alert without a named recipient is practically worthless, and what the logs must contain so that an incident is still traceable weeks later.
Key takeaways
- Availability monitoring only answers whether the technology responds; the expensive failures in mid-size companies happen precisely where every technical indicator stays green and yet no data moves between two systems any more.
- Sound monitoring works on four levels — availability, run, business check and outcome — and only the third level asks whether orders actually arrived today and whether the volumes on both sides match.
- Thresholds built from fixed numbers produce false alarms; a threshold becomes usable only when it combines three parts: an expected value from your own history, a tolerance band and a minimum duration before anything is reported.
- Every check needs a named person, a named deputy and a deadline after which the next escalation stage takes over; an alert without a named recipient is not monitoring but documentation of damage that has already occurred.
- Logs must record timestamp, the quantity checked, the expected value, the actual value, the assessment and the recipient, with a defined retention period — otherwise the very period under investigation is missing when an incident is reviewed.
What availability monitoring does not answer
Availability monitoring checks whether a system responds. At short intervals it queries a server, a service or an address and reports when no answer comes back. That is useful and belongs in every company's basic equipment. But this kind of check describes the state of the technology, not the result of the work. It cannot tell the difference between a transfer that moved twenty thousand line items and one that moved nothing — in both cases the server answered.
The failures that genuinely cost money in mid-size companies tend to look very similar: the expired access key nobody renewed. The additional mandatory field after an update that makes the target system reject every second record. The filter that kicks in after a range change and excludes an entire product group. The order status that was renamed and therefore no longer matches the rule. In all four cases the server runs continuously, the service is reachable, the backup is green. And still nothing arrives in business terms.
The difference between the two kinds of monitoring is organisational rather than technical: availability is defined by IT, business relevance is defined by the department. Nobody in IT knows what volume is normal on a Tuesday morning — the person who works with the orders every day knows that. Monitoring designed without that person inevitably checks only what can be measured technically.
A nightly run that transfers nothing
Four levels: from availability to outcome
It helps to think of monitoring not as one thing but as four levels building on each other. Each level answers its own question, each catches its own type of failure, and none replaces another. Leaving a level out creates a blind spot — not a problem as long as you know about it, dangerous when you believe it is covered.
Effort rises from level to level, but so does the benefit for day-to-day operations. The first two levels usually come with the technology or can be set up within a few hours. The third and fourth levels have to be described by someone from the business side, because that is where operational expectations are put into words. That is exactly why they are missing in many companies — not because of the technology, but because nobody writes the expectation down.
Level 1: Availability
Does the server respond, is the service running, is there enough disk space, is the certificate valid? These checks run every few minutes and report hard outages. They are quick to set up and cover the smaller share of the failures that actually occur.
Level 2: Run
Did the scheduled transfer run start, did it finish within the usual duration, were there aborted jobs? This level catches hanging processes and failed schedulers — but says nothing about whether any content was moved.
Level 3: Business check
Were orders transferred today, does the count match expectations, do documents exist on both sides with the same totals, are transactions sitting unusually long in the error area? This is where monitoring starts to reflect the business rather than the technology.
Level 4: Outcome
Do the stock figures in both systems agree at the end of the day, how many transactions had to be reworked manually, what share of handovers failed per week? This level does not assess individual runs but the quality of the connection over time.
In most companies whose support we take over, level 1 is in place, level 2 partly, level 3 rarely and level 4 almost never (project experience). Yet the fourth level is the only one that supports decisions: it shows which connection permanently creates rework and where a rebuild pays off. Which figures are suitable for that and how they are produced regularly is described under metrics and reporting.
The three questions a business check asks
A business check is not a complicated tool but a question that is asked and answered automatically every day. Three questions are enough to begin with, and they can be formulated for practically any connection regardless of which systems are involved: did anything arrive? Is the volume plausible? Is anything stuck?
The first question catches the complete standstill, the second the gradual gap, the third the backlog. Together they cover most of what goes wrong in practice. What matters is that the check runs outside the interface and pulls figures from both systems involved. A component that checks itself will report nothing precisely when it has failed.
- Number of transactions transferred per direction and day, compared with the number in the source system
- Total document values on both sides for the same period, compared down to the cent
- Age of the oldest unprocessed entry in the queue, measured in hours
- Number of records rejected by the target system, including the most frequent reason
- Timestamp of the last successful transfer per direction, independent of the run status
- Number of master data records that exist on one side and are missing on the other
These checks produce a short daily report that is sent even when everything is fine. That sounds like unnecessary mail but serves a practical purpose: a report that fails to arrive is itself an alert. Anyone notified only on errors cannot tell a quiet day from a monitoring system that has stopped working.
A report like this can be read in a few minutes and needs no technical background. That is precisely the point: the person reading it has to be able to judge whether 148 orders are plausible for that day. No tool can take over that judgement, and it is the reason business-level monitoring belongs in the department rather than solely in IT.
Thresholds: when a deviation becomes an incident
Once the checks are in place, the harder question follows: at what point should an alert be raised? Fixed numbers fail regularly. Setting a rule that fewer than fifty orders a day triggers an alert produces a message every Saturday, on every bridging day and during every shutdown period — and the check is switched off after three weeks. Setting the threshold so low that it never triggers leaves a check that is nothing but paperwork.
A usable threshold has three components. First an expected value derived from your own history, for example the average of the same weekday over the past eight weeks. Second a tolerance band around that value reflecting normal fluctuation. Third a minimum duration: a deviation only becomes an incident once it persists for a defined period — around thirty minutes for a transfer running every few minutes, a single missed run for a nightly job.
| Check | Threshold that works | Threshold that creates false alarms |
|---|---|---|
| Number of orders per day | Deviation beyond the tolerance band of the same weekday over eight weeks | Fixed lower limit regardless of weekday and season |
| Document totals on both sides | Difference other than zero, checked after the daily run has finished | Checking during the running reconciliation, where interim states are normal |
| Age of the oldest queued item | Older than the agreed processing deadline for the connection | Every entry waiting longer than a minute |
| Rejected records | Share above the usual level, or a new and previously unknown reason | Every single rejection as a separate alert |
| Timestamp of the last transfer | Longer ago than the agreed interval plus a grace period | Exactly the interval with no grace period, so every delay reports |
| Master data differences | New differences compared with the previous day | Total number of all differences, including known legacy cases |
In practice it works well to start wide and tighten the thresholds after two to four weeks, once the real fluctuation is known. That includes an operating calendar: public holidays, shutdown periods, stocktaking days and announced maintenance windows belong in the configuration, otherwise monitoring reports on exactly the days when nobody is available anyway. A threshold that has never triggered is not a quiet connection but an unverified assumption.
An alert with no owner is merely documentation
The most common reason monitoring goes nowhere in mid-size companies is not missing technology but missing ownership. The alert goes to a shared mailbox where everyone reads along and therefore nobody acts, or to a distribution list created three years ago that still contains people who left long ago. The failure has been detected technically — operationally nothing has happened.
Ownership means a named person, a named deputy, a deadline for the first response and a second stage that takes over when that deadline passes. This is not bureaucracy but the only way to turn an alert into an action. For most connections in a mid-size company three or four stages are enough, and they fit on half a page.
Stage 1: Alert to the responsible person
The alert goes to the person responsible in the department, with a subject line, the affected connection, the timestamp and a statement of which value lies outside expectations. The response deadline during business hours is typically two hours, for nightly runs the following morning.
Stage 2: Deputy and technical support
If no response comes back, the same alert goes to the deputy and to technical support. What matters here is that the alert is not merely repeated but clearly marked as an escalation, including the time already elapsed.
Stage 3: Inform management
After a defined total duration — often half a working day — management is informed, because from that point operational decisions are due: switch to manual entry, inform customers, hold deliveries? Those decisions do not belong to the technical side.
Stage 4: All-clear and a short note
Once resolved, an all-clear goes to everyone previously informed so that nobody intervenes twice. Three lines go into the log: cause, how it was detected, time until the alert. That note later forms the basis for deciding which connection needs attention.
The alert goes to the person responsible in the department, with a subject line, the affected connection, the timestamp and a statement of which value lies outside expectations. The response deadline during business hours is typically two hours, for nightly runs the following morning.
If no response comes back, the same alert goes to the deputy and to technical support. What matters here is that the alert is not merely repeated but clearly marked as an escalation, including the time already elapsed.
After a defined total duration — often half a working day — management is informed, because from that point operational decisions are due: switch to manual entry, inform customers, hold deliveries? Those decisions do not belong to the technical side.
Once resolved, an all-clear goes to everyone previously informed so that nobody intervenes twice. Three lines go into the log: cause, how it was detected, time until the alert. That note later forms the basis for deciding which connection needs attention.
Where support is outsourced, the escalation chain belongs in the agreement: who reports, who responds, within which times, and what applies outside business hours. Without that, the classic gap opens up in which both sides assume the other is taking care of it. What such a support arrangement can look like is described under IT operations.
Monitoring is only as good as the name behind each check — not as good as the technology that runs it.
Logs: what gets recorded and how long it stays
Logs are the part that seems least important during setup and is missed most when an incident is reviewed. The question is never whether anything is logged at all — almost every system writes something. The question is whether the log contains what you will need later, and whether it still exists at that point.
For a business check, six pieces of information belong in every line: timestamp, the quantity checked, the expected value, the actual value, the assessment and the recipient of the alert. Without the expected value it is impossible to say afterwards whether the check should have triggered at all. Without the recipient there is no way to trace who was informed. Both are regularly omitted because they look redundant in normal operation.
timestamp : 2026-06-11 18:45:03
check : order_count_day
connection : Inventory system -> Accounting system
expected : 132 to 176 (average Thu, 8 weeks)
actual : 148
assessment : within tolerance band
recipient : no alert raised
retain_until : 2027-06-11The second half concerns retention. In many environments logs overwrite themselves after a few days because that is the default setting. If an error only surfaces during the monthly close, exactly the period in question has disappeared. What works is an explicitly defined retention period per connection, separated into technical logs and business check logs — technical lines may live shorter, business lines should outlast at least one completed closing cycle.
Retention and data protection have to be considered together
Where these decisions are recorded is not a matter of taste: they belong in the description of the connection, together with direction, interval, ownership and restart procedure. Monitoring that is only configured in tools and described nowhere is lost with the next system change. How to keep such descriptions short instead of turning them into a binder is shown under process documentation.
False alarms cost more than missing alarms
There is one state worse than no monitoring: monitoring nobody believes any more. It develops along the same path every time. At the start the thresholds are too tight, several alerts arrive every day and most of them require no action. After two weeks the recipient sets up a rule that files those messages in a folder. From that moment monitoring is effectively switched off, even though every report states that it is active.
There are several remedies and they cost little. Bundle alerts instead of sending them individually: one message per connection and incident, not per failed record. Suppress repetitions while an incident is open and send reminders at fixed intervals instead. Mute known, planned states — but only for a limited time and with automatic reactivation, because a permanently muted check is a deleted check, only less honest.
The figure that shows the state of your monitoring
The second test is the deliberate one. Monitoring that has never triggered is an assumption. Once a quarter a check is made to fire on purpose — invalidate access credentials temporarily, stop the service briefly, feed in a test record with a wrong quantity. What is measured is not only whether an alert arrives, but when it arrives and who receives it. In practice the recipient list fails more often than the technology (project experience).
Rolling it out in manageable steps
The usual mistake during rollout is insisting on completeness. Anyone wanting to monitor every connection on four levels at once needs a project, a budget and coordination — and therefore never starts. The reverse order works better: begin with the connection whose three-day outage would create the largest backlog, and set up the simplest effective check there.
The following order can be worked through alongside daily business. Each step delivers value on its own, and you can stop after any step without leaving a half-finished state. The organisational part dominates: points one to three are mostly agreement, real implementation effort only starts at point four.
- List the connections: source, target, direction, interval, responsible person in the department, technical support. Without this list every further measure remains piecemeal.
- Define ownership and deadlines per connection, including the deputy and what applies outside business hours.
- Formulate one business question per connection, in the language of the department: what has to have arrived by when for the day to count as normal?
- Set up the simplest check for it, usually a volume comparison once a day, and send the report even when the run was clean.
- Tighten the thresholds after two to four weeks once the real fluctuation is known, and add the operating calendar.
- Define log content and retention periods and record them in the description of the connection.
- Run a deliberate trigger test once a quarter and check whether the recipient list is still correct.
For companies without their own IT department this still means noticeable effort, especially in points four and six. It can be handled as a one-off setup followed by ongoing support; recording all connections with direction, interval and ownership is part of a process analysis, and how setup and ongoing support are charged is set out under pricing. The decisive point remains the same regardless of who implements it: monitoring without named ownership is a technical indicator. Only with a name, a deadline and an escalation path does it become part of operations.
Related Articles
When an interface fails: spotting silent outages
Detecting, reporting and bridging silent interface failures: heartbeat, time windows, volume reconciliation, queueing and a defined restart procedure at work.
Automating file imports: six failure modes and their fixes
Daily file transfers between two systems: six typical failure modes from delimiters to duplicate deliveries, and how to check, log and report every one of them.
Invoice checks automated: order, goods receipt, invoice
Matching order, goods receipt and invoice by machine: which fields are compared, where the tolerance band sits, who owns the exception and what stays manual.