Most interface problems give no warning. There is no alarm, no red message on screen and no call from the vendor. The transfer between two systems simply stops, and the business carries on as if nothing had happened: orders are taken, delivery notes printed, invoices written. Just no longer on both sides. Anyone running an interface therefore needs to think less about the normal case than about the outage. How is it noticed? Who finds out? What happens to the data in the meantime? And how does the business get back to a clean state afterwards? This article describes the five building blocks that make a connection fit for daily operation, and closes by showing how to check an existing interface within minutes to see whether anyone is watching at all.
Key takeaways
- The dangerous outage is the silent one: the transfer stops, neither system reports anything, and the fault only surfaces when a customer asks about a delivery or accounting finds a gap in the figures.
- Detection needs three patterns side by side: a heartbeat that regularly confirms the connection is alive, a time window by which a transfer must have arrived, and a volume reconciliation comparing record counts on both sides.
- An alert only counts as an alert once it reaches a named responsible person, identifies the affected transaction and contains an instruction to act, instead of landing in a shared mailbox nobody opens.
- A queue prevents data loss: undelivered records stay put with a timestamp, a status and a failure reason until the restart procedure replays them in a defined order and without creating double postings.
- Whether an existing interface is monitored comes down to three questions: who receives the alert, where are the logs, and what happens during a deliberately triggered test outage — if one of them stays unanswered, the connection is unwatched.
The loud outage is the harmless one
When a server is unreachable, everybody notices within minutes. Nobody can log in, the phone rings, the matter gets dealt with. Such incidents are unpleasant, but they report themselves: the damage is limited to the downtime, and there is nothing to catch up on afterwards, because nothing was created during that period anyway. The commercial impact usually stays manageable because the response starts immediately.
An interface behaves differently. It works in the background, often overnight, often every few minutes, and nobody sits in front of it. If it fails, the user interface looks exactly the same. Sales keeps entering orders, the warehouse keeps posting stock movements, accounting keeps issuing invoices. Only the reconciliation between the systems no longer happens. Every day without a transfer creates additional transactions that later have to be replayed by hand — and in the correct order, because stock levels, prices and payment states build on one another.
In practice, discovery rarely takes hours; it takes days. And the outage is usually not found by IT but by a customer missing an order confirmation, by a technician standing in front of an empty storage bin, or by accounting during the monthly close. Until then the number of affected transactions is unknown, and that uncertainty is the real cost driver: nobody can say whether twelve or twelve hundred records are missing, so everything has to be checked.
What a silent failure is
Three patterns for noticing a failure at all
Detection is not a question of expensive tooling but of asking the right question. An interface can fail in three fundamentally different ways: the connection itself is dead, the connection is alive but nothing arrives, or something arrives but incompletely. Each of these has its own check pattern. Using only one of them covers only part of the possible failures.
The three patterns complement each other and can be implemented in almost any environment — even where the systems involved bring no monitoring functions of their own. What matters is not the technology but the fact that the check happens outside the interface. A component that is supposed to report its own failure can no longer do so once it has failed.
Heartbeat
The interface reports at a fixed rhythm that it is running — a short entry or call every few minutes, for example. If that sign of life stops, the watching component raises an alert. A heartbeat catches dead processes and crashed services, but says nothing about whether data is actually flowing.
Time window
For every schedulable transfer, a deadline is defined: the overnight stock list by seven in the morning, the invoice handover by end of day. If nothing arrives by then, that counts as an incident. This pattern also catches the case where the connection is technically alive but nobody is sending anything.
Volume reconciliation
Both sides count what they processed within a period, and the numbers are compared. If they differ, records are missing or were processed twice. Volume reconciliation is the only pattern that finds partial failures — for instance when only certain transactions are silently discarded.
The effort for these three checks is modest if it is considered while the interface is being built. Added afterwards, volume reconciliation is usually the most expensive part, because both systems have to provide a reliable count. A pragmatic middle step: introduce the heartbeat and the time window first, run the volume reconciliation as a daily report reviewed by eye, and automate it only once it is clear which deviations occur in normal operation anyway.
The alert has to reach a human being
Detection without notification is worthless. Many companies do have a log in which the error is neatly recorded — but nobody reads it, because it sits in a directory on a server that somebody would have to open deliberately. An incident that exists only in the log has effectively not been detected. The alert has to reach the person; the person should not have to go looking for the alert.
The opposite matters just as much: too many alerts are worse than none. If an interface sends a message for every single failed attempt, a bad day produces hundreds of messages. Before long the recipient sets up a rule that files them into a folder — and from that moment on, monitoring is effectively switched off even though it works technically. A workable pattern is therefore: one message when the incident starts, one reminder after a defined period, one message when it is resolved. Everything else belongs in the log, not in the inbox.
- A recipient named as a person plus a named deputy, not a shared address such as info or it
- A subject line that shows, without opening, which connection is affected and since when
- The affected transaction or period, so the scale can be assessed straight away
- One concrete first instruction that a deputy without deep knowledge can carry out
- A second channel in case the mail system itself is affected, for instance a text message
- An all-clear message after resolution, so nobody intervenes twice
If the connections are looked after externally, the contract should state who notifies whom in an incident and within what time a response follows. Without that, the classic gap appears: the service provider sees the alert and assumes it is an internal matter, while the company assumes the provider is dealing with it. What such a support arrangement looks like is described under IT operations.
A queue instead of data loss
The second big difference between a robust and an improvised interface lies in what happens to the data while the other side is unreachable. A simply built connection tries the handover exactly once. If it fails, the record is gone — not deleted, but not transferred either, and nobody knows which ones were affected. The follow-up work then consists of pulling lists from both systems and holding them against each other.
A queue turns that around. Every transaction to be transferred is written to intermediate storage first and marked as done only after the other side confirms receipt. If the handover fails, the entry stays put with a timestamp, a failure reason and an attempt counter. Repeat attempts follow at growing intervals, so an overloaded target system is not put under additional strain. After a defined number of attempts, the entry moves into an area reserved for cases that need a human decision.
id : 20260608-000417
created : 2026-06-08 07:42:11
direction : inventory system -> accounting system
transaction : invoice R-2026-4417
status : waiting
attempts : 3
last_error : target system not responding (timeout)
retry_after : 2026-06-08 08:12:11
checksum : e1f9c2a4How repeat attempts are handled is crucial. If the target system has already accepted a transaction and only the confirmation was lost, another attempt would create a double posting. This is prevented by a unique identifier per transaction that the target system checks before processing. In plain terms: a second call carrying the same identifier changes nothing. For invoices, payments and stock movements this is not a subtlety but the precondition for a restart being safe at all.
Restart: planned, not improvised
The moment the connection comes back is the most dangerous of the whole incident. Several thousand transactions may now be sitting in the queue, and if they are fired off unthrottled and in arbitrary order, new problems appear: the target system buckles under the load, stock postings run in the wrong sequence, price changes take effect too late. Acting without a defined procedure at that moment often produces more follow-up work than the outage itself caused.
So the restart procedure belongs in writing before it is needed — short, one page, in the language of the people expected to carry it out. It also has to work when the person who built the interface happens to be unreachable. That is the real test: can the holiday cover act on the description without phoning anyone?
Step 1: Confirm the cause, do not assume it
Before the restart, check that the fault is genuinely fixed: a single test transfer with one uncritical record. If it goes through, the way is clear. Opening the queue while the cause persists only increases the attempt counter and prolongs the outage.
Step 2: Define the order
Transactions that build on one another are replayed in their original sequence — master data first, then stock, then documents. Within a type, sort by creation time. That order belongs in the written procedure; it is not invented during the incident.
Step 3: Work through the queue at a throttled pace
The queue is drained at a limited rate so the target system stays responsive for day-to-day work. With a large backlog it is worth starting outside core hours. Progress stays visible so it is clear whether the backlog is actually shrinking.
Step 4: Run the reconciliation
Once the queue is empty, repeat the volume reconciliation for the affected period. If the counts match on both sides, the technical part is complete. Deviations are listed individually rather than estimated in round numbers.
Step 5: Sign off and record
A named person declares the restart complete and records the time, the duration, the affected transactions and the cause. Only then does the incident count as over — including for the departments that have been working with restrictions until that point.
Before the restart, check that the fault is genuinely fixed: a single test transfer with one uncritical record. If it goes through, the way is clear. Opening the queue while the cause persists only increases the attempt counter and prolongs the outage.
Transactions that build on one another are replayed in their original sequence — master data first, then stock, then documents. Within a type, sort by creation time. That order belongs in the written procedure; it is not invented during the incident.
The queue is drained at a limited rate so the target system stays responsive for day-to-day work. With a large backlog it is worth starting outside core hours. Progress stays visible so it is clear whether the backlog is actually shrinking.
Once the queue is empty, repeat the volume reconciliation for the affected period. If the counts match on both sides, the technical part is complete. Deviations are listed individually rather than estimated in round numbers.
A named person declares the restart complete and records the time, the duration, the affected transactions and the cause. Only then does the incident count as over — including for the departments that have been working with restrictions until that point.
Follow-up work: closing the gap and describing it
The technical restart rarely ends the matter. During the incident, business decisions were made on false assumptions: something was sold that had long been reserved in the other system, a reminder went to a customer whose payment simply had not been transferred, an order was created twice because the first one never appeared elsewhere. No restart resolves these cases; they need a review by the people who know the business.
A short catch-up list works well: which transactions were created during the outage window, which of them were communicated externally, and where is a correction needed? Responsibility sits with the department, not with IT — IT supplies the list, the assessment happens in the business. Two or three categories are enough: uncritical, correct internally, clarify with the customer.
An incident is also a source of data
The procedure belongs in the interface description: what is transferred, in which direction, at what rhythm, who acts during an incident and how the restart runs. Where transfers are relevant to bookkeeping, this also touches documentation requirements under national accounting rules; assessing the individual case belongs with tax advisers. How to build such a description without producing a ring binder is shown under process documentation.
An interface is only finished once there is a written answer to what happens while it is not working.
How to tell whether your interface is monitored
Most connections in mid-size companies have grown over the years. Whoever built them is sometimes no longer with the company. Whether monitoring exists can still be established without technical knowledge — with questions that any responsible person must be able to answer. The attitude matters: this is not about blame but about the current state. Many interfaces run inconspicuously for years and were therefore never a topic.
You can walk through the table below with your internal IT team or the external provider. If a row produces no solid answer, that is not proof of a problem, but it is a point for the list. If three or more rows stay open, the connection is running unwatched.
| Question to ask | Solid answer | Warning sign |
|---|---|---|
| Who receives an alert if the transfer does not arrive? | The name of a person and a deputy | It gets written to the log |
| When was the last incident noticed, and how? | Date and detection route are known | There has never been an incident |
| Where are the logs kept and for how long? | Location and retention period are defined | Logs overwrite themselves daily |
| What happens to data when the other side does not answer? | It waits in a queue | The attempt is discarded |
| How does the restart run after a longer outage? | There is a written sequence | The colleague who knows it does that |
| When was the monitoring last triggered deliberately? | A test outage within recent months | Never tested |
The last row is the most telling. Monitoring that has never fired is an assumption. The test is simple and takes about fifteen minutes: temporarily invalidate the credentials or stop the service on purpose, then wait to see whether and when an alert arrives and who receives it. Pick a period with low volume and inform the affected departments briefly in advance. A complete inventory of all connections with direction, rhythm and responsibility is part of a process analysis.
The order in which to retrofit all this
Introducing all five building blocks at once overwhelms most companies — and it is not necessary either. The benefit is unevenly distributed: the first steps cost little and cover the bulk of cases, the later ones take more effort and pay off only for critical connections. If you have to prioritise, go by potential damage: which connection would create the biggest backlog after a three-day outage?
The following order has proven itself. It can be worked through step by step, each step delivers value on its own, and you can stop after any step without leaving a half-finished state behind.
- Take inventory: list every connection with source, target, direction, rhythm and responsible person. Without this list, every further measure is piecemeal.
- Set up the alerting route: for each connection define a named person and a deputy, plus the subject line and content of the incident message.
- Monitor time windows: define a deadline for every schedulable transfer and raise an alert if nothing has arrived. The fastest benefit for the least effort.
- Add a heartbeat: have running connections send a sign of life and watch it from outside, so crashed services are noticed too.
- Introduce a queue: start with the connections where data loss does the most damage, usually documents and payments.
- Write down the restart: one page per connection, written for the deputy, and rehearsed once in a test case.
- Automate volume reconciliation: last, because it delivers the most reliable figures but also needs the most coordination between systems.
The effort stays contained if points one to three come first: they are largely organisation, not programming. Real implementation effort starts at point four. An unwatched interface is not a technical risk but a commercial one — the damage arises not in the technology but in the follow-up work and in customer trust. How review, retrofitting and ongoing support are charged is set out under pricing.
The same order applies to new builds
Related Articles
Monitoring interfaces properly: beyond the server being up
Monitoring interfaces beyond availability: business-level checks, thresholds without false alarms, escalation to a named owner and logs with defined retention.
Automating file imports: six failure modes and their fixes
Daily file transfers between two systems: six typical failure modes from delimiters to duplicate deliveries, and how to check, log and report every one of them.
Webhooks instead of polling: when the switch pays off
Poll on a schedule or get notified? What each approach costs in daily operation, what a receiver must provide and when polling remains the safer choice.