Skip to content
Automation & interfaces

Monitoring interfaces properly: beyond the server being up

Monitoring interfaces beyond availability: business-level checks, thresholds without false alarms, escalation to a named owner and logs with defined retention.

13 min read ÜberwachungSchnittstellenEskalationProtokollierungSchwellwerte

Most companies do have monitoring in place — it simply answers the wrong question. One service reports that the server is reachable, a second that disk space is sufficient, a third that the backup completed. All three can show green while not a single order has moved from the inventory system to the accounting system for four days. None of those checks looks inside the transaction. Anyone running an interface, or having one run for them, needs monitoring that knows the business rather than only the technology: did any orders arrive today? Do the volumes match what the other side sent? Is anything stuck that was flowing yesterday? This article describes how to build such business-level checks, how to set thresholds without drowning in false alarms, why an alert without a named recipient is practically worthless, and what the logs must contain so that an incident is still traceable weeks later.

Key takeaways

  • Availability monitoring only answers whether the technology responds; the expensive failures in mid-size companies happen precisely where every technical indicator stays green and yet no data moves between two systems any more.
  • Sound monitoring works on four levels — availability, run, business check and outcome — and only the third level asks whether orders actually arrived today and whether the volumes on both sides match.
  • Thresholds built from fixed numbers produce false alarms; a threshold becomes usable only when it combines three parts: an expected value from your own history, a tolerance band and a minimum duration before anything is reported.
  • Every check needs a named person, a named deputy and a deadline after which the next escalation stage takes over; an alert without a named recipient is not monitoring but documentation of damage that has already occurred.
  • Logs must record timestamp, the quantity checked, the expected value, the actual value, the assessment and the recipient, with a defined retention period — otherwise the very period under investigation is missing when an incident is reviewed.

What availability monitoring does not answer

Availability monitoring checks whether a system responds. At short intervals it queries a server, a service or an address and reports when no answer comes back. That is useful and belongs in every company's basic equipment. But this kind of check describes the state of the technology, not the result of the work. It cannot tell the difference between a transfer that moved twenty thousand line items and one that moved nothing — in both cases the server answered.

The failures that genuinely cost money in mid-size companies tend to look very similar: the expired access key nobody renewed. The additional mandatory field after an update that makes the target system reject every second record. The filter that kicks in after a range change and excludes an entire product group. The order status that was renamed and therefore no longer matches the rule. In all four cases the server runs continuously, the service is reachable, the backup is green. And still nothing arrives in business terms.

The difference between the two kinds of monitoring is organisational rather than technical: availability is defined by IT, business relevance is defined by the department. Nobody in IT knows what volume is normal on a Tuesday morning — the person who works with the orders every day knows that. Monitoring designed without that person inevitably checks only what can be measured technically.

A nightly run that transfers nothing

The recurring case in practice: the nightly transfer run starts on schedule at 2 a.m., processes zero records, finishes without an error and writes an entry with the status successful. Availability monitoring is green, run monitoring is green, the morning report is green. Only a business check — comparing the number transferred with the number of open transactions in the source — would have triggered that same night.

Four levels: from availability to outcome

It helps to think of monitoring not as one thing but as four levels building on each other. Each level answers its own question, each catches its own type of failure, and none replaces another. Leaving a level out creates a blind spot — not a problem as long as you know about it, dangerous when you believe it is covered.

Effort rises from level to level, but so does the benefit for day-to-day operations. The first two levels usually come with the technology or can be set up within a few hours. The third and fourth levels have to be described by someone from the business side, because that is where operational expectations are put into words. That is exactly why they are missing in many companies — not because of the technology, but because nobody writes the expectation down.

Level 1: Availability

Does the server respond, is the service running, is there enough disk space, is the certificate valid? These checks run every few minutes and report hard outages. They are quick to set up and cover the smaller share of the failures that actually occur.

Level 2: Run

Did the scheduled transfer run start, did it finish within the usual duration, were there aborted jobs? This level catches hanging processes and failed schedulers — but says nothing about whether any content was moved.

Level 3: Business check

Were orders transferred today, does the count match expectations, do documents exist on both sides with the same totals, are transactions sitting unusually long in the error area? This is where monitoring starts to reflect the business rather than the technology.

Level 4: Outcome

Do the stock figures in both systems agree at the end of the day, how many transactions had to be reworked manually, what share of handovers failed per week? This level does not assess individual runs but the quality of the connection over time.

In most companies whose support we take over, level 1 is in place, level 2 partly, level 3 rarely and level 4 almost never (project experience). Yet the fourth level is the only one that supports decisions: it shows which connection permanently creates rework and where a rebuild pays off. Which figures are suitable for that and how they are produced regularly is described under metrics and reporting.

The three questions a business check asks

A business check is not a complicated tool but a question that is asked and answered automatically every day. Three questions are enough to begin with, and they can be formulated for practically any connection regardless of which systems are involved: did anything arrive? Is the volume plausible? Is anything stuck?

The first question catches the complete standstill, the second the gradual gap, the third the backlog. Together they cover most of what goes wrong in practice. What matters is that the check runs outside the interface and pulls figures from both systems involved. A component that checks itself will report nothing precisely when it has failed.

  • Number of transactions transferred per direction and day, compared with the number in the source system
  • Total document values on both sides for the same period, compared down to the cent
  • Age of the oldest unprocessed entry in the queue, measured in hours
  • Number of records rejected by the target system, including the most frequent reason
  • Timestamp of the last successful transfer per direction, independent of the run status
  • Number of master data records that exist on one side and are missing on the other

These checks produce a short daily report that is sent even when everything is fine. That sounds like unnecessary mail but serves a practical purpose: a report that fails to arrive is itself an alert. Anyone notified only on errors cannot tell a quiet day from a monitoring system that has stopped working.

Terminal
$ daily-report interface-inventory-accounting --date 2026-06-11
Orders in source : 148 Orders transferred : 148 OK Document total source : 214,880.42 EUR Document total target : 214,880.42 EUR OK Oldest queued item : 0 h 12 min OK Rejected records : 3 reason: customer number unknown Last transfer : 2026-06-11 18:40 OK Master data difference : 12 items missing in target CHECK

A report like this can be read in a few minutes and needs no technical background. That is precisely the point: the person reading it has to be able to judge whether 148 orders are plausible for that day. No tool can take over that judgement, and it is the reason business-level monitoring belongs in the department rather than solely in IT.

Thresholds: when a deviation becomes an incident

Once the checks are in place, the harder question follows: at what point should an alert be raised? Fixed numbers fail regularly. Setting a rule that fewer than fifty orders a day triggers an alert produces a message every Saturday, on every bridging day and during every shutdown period — and the check is switched off after three weeks. Setting the threshold so low that it never triggers leaves a check that is nothing but paperwork.

A usable threshold has three components. First an expected value derived from your own history, for example the average of the same weekday over the past eight weeks. Second a tolerance band around that value reflecting normal fluctuation. Third a minimum duration: a deviation only becomes an incident once it persists for a defined period — around thirty minutes for a transfer running every few minutes, a single missed run for a nightly job.

CheckThreshold that worksThreshold that creates false alarms
Number of orders per dayDeviation beyond the tolerance band of the same weekday over eight weeksFixed lower limit regardless of weekday and season
Document totals on both sidesDifference other than zero, checked after the daily run has finishedChecking during the running reconciliation, where interim states are normal
Age of the oldest queued itemOlder than the agreed processing deadline for the connectionEvery entry waiting longer than a minute
Rejected recordsShare above the usual level, or a new and previously unknown reasonEvery single rejection as a separate alert
Timestamp of the last transferLonger ago than the agreed interval plus a grace periodExactly the interval with no grace period, so every delay reports
Master data differencesNew differences compared with the previous dayTotal number of all differences, including known legacy cases

In practice it works well to start wide and tighten the thresholds after two to four weeks, once the real fluctuation is known. That includes an operating calendar: public holidays, shutdown periods, stocktaking days and announced maintenance windows belong in the configuration, otherwise monitoring reports on exactly the days when nobody is available anyway. A threshold that has never triggered is not a quiet connection but an unverified assumption.

An alert with no owner is merely documentation

The most common reason monitoring goes nowhere in mid-size companies is not missing technology but missing ownership. The alert goes to a shared mailbox where everyone reads along and therefore nobody acts, or to a distribution list created three years ago that still contains people who left long ago. The failure has been detected technically — operationally nothing has happened.

Ownership means a named person, a named deputy, a deadline for the first response and a second stage that takes over when that deadline passes. This is not bureaucracy but the only way to turn an alert into an action. For most connections in a mid-size company three or four stages are enough, and they fit on half a page.

The alert goes to the person responsible in the department, with a subject line, the affected connection, the timestamp and a statement of which value lies outside expectations. The response deadline during business hours is typically two hours, for nightly runs the following morning.

Where support is outsourced, the escalation chain belongs in the agreement: who reports, who responds, within which times, and what applies outside business hours. Without that, the classic gap opens up in which both sides assume the other is taking care of it. What such a support arrangement can look like is described under IT operations.

Monitoring is only as good as the name behind each check — not as good as the technology that runs it.

Project experience

Logs: what gets recorded and how long it stays

Logs are the part that seems least important during setup and is missed most when an incident is reviewed. The question is never whether anything is logged at all — almost every system writes something. The question is whether the log contains what you will need later, and whether it still exists at that point.

For a business check, six pieces of information belong in every line: timestamp, the quantity checked, the expected value, the actual value, the assessment and the recipient of the alert. Without the expected value it is impossible to say afterwards whether the check should have triggered at all. Without the recipient there is no way to trace who was informed. Both are regularly omitted because they look redundant in normal operation.

Log line of a business check (schema)
timestamp        : 2026-06-11 18:45:03
check            : order_count_day
connection       : Inventory system -> Accounting system
expected         : 132 to 176 (average Thu, 8 weeks)
actual           : 148
assessment       : within tolerance band
recipient        : no alert raised
retain_until     : 2027-06-11

The second half concerns retention. In many environments logs overwrite themselves after a few days because that is the default setting. If an error only surfaces during the monthly close, exactly the period in question has disappeared. What works is an explicitly defined retention period per connection, separated into technical logs and business check logs — technical lines may live shorter, business lines should outlast at least one completed closing cycle.

Retention and data protection have to be considered together

Logs covering transfers relevant to accounting touch requirements for process documentation and traceable records. At the same time they often contain personal data — processor names, customer numbers, sometimes content in plain text inside error messages. The two pull in opposite directions: retain and delete. Retention periods should therefore be defined per log type and, in the individual case, agreed with your tax and legal advisers; where employee behaviour or performance could be derived, works council codetermination applies as well. This article does not replace legal advice.

Where these decisions are recorded is not a matter of taste: they belong in the description of the connection, together with direction, interval, ownership and restart procedure. Monitoring that is only configured in tools and described nowhere is lost with the next system change. How to keep such descriptions short instead of turning them into a binder is shown under process documentation.

False alarms cost more than missing alarms

There is one state worse than no monitoring: monitoring nobody believes any more. It develops along the same path every time. At the start the thresholds are too tight, several alerts arrive every day and most of them require no action. After two weeks the recipient sets up a rule that files those messages in a folder. From that moment monitoring is effectively switched off, even though every report states that it is active.

There are several remedies and they cost little. Bundle alerts instead of sending them individually: one message per connection and incident, not per failed record. Suppress repetitions while an incident is open and send reminders at fixed intervals instead. Mute known, planned states — but only for a limited time and with automatic reactivation, because a permanently muted check is a deleted check, only less honest.

The figure that shows the state of your monitoring

Once a quarter, count two numbers: how many alerts were raised, and in how many cases somebody actually did something as a result. If the share of actionable alerts stays permanently low, business is not quiet — the monitoring is set too loud. This review takes half an hour and settles any discussion about perceived alert fatigue.

The second test is the deliberate one. Monitoring that has never triggered is an assumption. Once a quarter a check is made to fire on purpose — invalidate access credentials temporarily, stop the service briefly, feed in a test record with a wrong quantity. What is measured is not only whether an alert arrives, but when it arrives and who receives it. In practice the recipient list fails more often than the technology (project experience).

Rolling it out in manageable steps

The usual mistake during rollout is insisting on completeness. Anyone wanting to monitor every connection on four levels at once needs a project, a budget and coordination — and therefore never starts. The reverse order works better: begin with the connection whose three-day outage would create the largest backlog, and set up the simplest effective check there.

The following order can be worked through alongside daily business. Each step delivers value on its own, and you can stop after any step without leaving a half-finished state. The organisational part dominates: points one to three are mostly agreement, real implementation effort only starts at point four.

  1. List the connections: source, target, direction, interval, responsible person in the department, technical support. Without this list every further measure remains piecemeal.
  2. Define ownership and deadlines per connection, including the deputy and what applies outside business hours.
  3. Formulate one business question per connection, in the language of the department: what has to have arrived by when for the day to count as normal?
  4. Set up the simplest check for it, usually a volume comparison once a day, and send the report even when the run was clean.
  5. Tighten the thresholds after two to four weeks once the real fluctuation is known, and add the operating calendar.
  6. Define log content and retention periods and record them in the description of the connection.
  7. Run a deliberate trigger test once a quarter and check whether the recipient list is still correct.

For companies without their own IT department this still means noticeable effort, especially in points four and six. It can be handled as a one-off setup followed by ongoing support; recording all connections with direction, interval and ownership is part of a process analysis, and how setup and ongoing support are charged is set out under pricing. The decisive point remains the same regardless of who implements it: monitoring without named ownership is a technical indicator. Only with a name, a deadline and an escalation path does it become part of operations.

This article is based on data from: the German federal information security authority (BSI), in particular its IT-Grundschutz recommendations on logging and detection, and our own project experience from interface and operations projects in mid-size companies.

Related Articles

Automation & interfaces

When an interface fails: spotting silent outages

Detecting, reporting and bridging silent interface failures: heartbeat, time windows, volume reconciliation, queueing and a defined restart procedure at work.

13 min read
Automation & interfaces

Automating file imports: six failure modes and their fixes

Daily file transfers between two systems: six typical failure modes from delimiters to duplicate deliveries, and how to check, log and report every one of them.

13 min read
Automation & interfaces

Invoice checks automated: order, goods receipt, invoice

Matching order, goods receipt and invoice by machine: which fields are compared, where the tolerance band sits, who owns the exception and what stays manual.

18 min read