Skip to content
Automation & interfaces

Automating file imports: six failure modes and their fixes

Daily file transfers between two systems: six typical failure modes from delimiters to duplicate deliveries, and how to check, log and report every one of them.

13 min read DateiimportCSVFehlerbehandlungProtokollierungSchnittstellen

In many companies part of the daily business hangs on a single file. The inventory system drops a stock list overnight, time recording delivers yesterday's hours in the morning, a supplier sends a price list once a week. As long as the file looks the way it did yesterday, nobody gives it a thought. If it fails to arrive or quietly changes its format, that often surfaces days later: in wrong stock figures, duplicate postings or a report that does not match anyone's expectation. This article describes the six failure modes that regularly bring a file import down, how each of them can be caught, and why a carefully built file transfer is often preferable to a hastily assembled API.

Key takeaways

  • A file import is an interface, even if nobody calls it one: it runs without operators, without a screen and without questions, which is why its failure normally goes unnoticed until downstream figures stop adding up.
  • Six failure modes account for the bulk of incidents: the wrong delimiter, a different character encoding, an ambiguous date format, missing or shifted columns, a duplicate delivery of the same file, and an empty delivery without data rows.
  • Every one of these can be detected before processing if the sending and receiving side share a written file agreement: field list, order, delimiter, encoding, date and number format, mandatory fields and the expected delivery volume.
  • Three storage folders, a checksum against double processing and one log entry per run turn a script into a traceable procedure that can still be reconstructed months later.
  • The most expensive case is the file that never arrives, because it produces no error at all: that is why you monitor an expected window and raise an alert when nothing has come in by the agreed time.

Why the daily file import is underestimated

When mid-size companies talk about interfaces, most people picture an API, credentials and a project with a specification. Everyday practice looks different. A substantial share of data exchange between systems runs on files: a list in a network folder, a location on a transfer server, an attachment in a shared mailbox. In Bitkom's surveys on digitisation, the lack of data exchange between systems regularly appears as an obstacle (Bitkom). The file is the quiet answer to that problem, and nobody ever ran it as a project.

Typically the routine was set up years ago because two systems needed to talk to each other and there was no time for anything more elaborate. Someone configured an export, someone else an import routine, and it has been running ever since. The knowledge about what each column means sits in the heads of one or two people. Documentation is rare, an agreement about the format usually absent. That is precisely what makes the routine fragile: it works as long as nothing changes on either side.

The decisive difference from a data entry screen is the missing feedback. Leave a mandatory field empty in the order system and you get a message. Deliver a file with the wrong delimiter and you get nothing. In the better case processing stops and nobody notices. In the worse case it does not stop but writes nonsense into the target system: stock levels at zero, prices off by a factor of a thousand, dates in the wrong month. An import without feedback is not an import without errors, it is an import without knowledge of errors.

The silent failure is the normal case

Reviewing incidents afterwards, the surprise is rarely the fault itself but the time it took to detect it (field experience). An import that delivered nothing for three weeks does not cost you three weeks of missing data, it costs you the rework on everything that was decided on outdated figures in the meantime: purchase orders, commitments to customers, reports.

Six failure modes that can hit any import

The variety of possible faults looks unmanageable at first, but it is not. In practice almost every incident traces back to one of six patterns. Anyone who knows these six can build a check chain that runs before the actual processing and, in case of doubt, rejects rather than half-processes. The overview below assigns to each pattern how it shows up in daily operation and what catches it.

Failure modeHow it shows upCountermeasure before processing
Wrong delimiterAll values end up in one column, or the row breaks in the middle of a textRead the delimiter from the file agreement, check the field count per row
Different character encodingAccented characters appear as replacement symbols, search no longer finds the articleFix the encoding, convert on import, report unknown characters
Ambiguous date formatThe fifth of June becomes the fifth of May, or a date turns into a numberDefine one binding format and validate every field against it
Missing or shifted columnsValues land in the wrong field, prices end up in the quantityCompare the header row with the field list, never assume the order
Duplicate deliveryPostings and movements appear twice, stock figures drift apartKeep a checksum and a run identifier, reject files already processed
Empty or truncated fileOnly the header row arrives, a stock level is set to zeroDefine a minimum volume per delivery, report large deviations from the previous day

One remark on sequence: these checks belong before processing, not in the middle of it. A file that was half imported and then failed leaves behind a state nobody can describe cleanly. Check first, then take it over completely or reject it completely. That is the most important design rule, and also the one most often missing from routines that simply grew over time.

Delimiters and encoding: the two quiet destroyers

A text file with separated values looks simple but is not a uniformly defined format. There are semicolons and commas as delimiters, quotation marks as text qualifiers, different rules for quotes inside a text and several conventions for line breaks. If the sending system changes its default after an update, the file is still formally a text file, but for the receiving side it is a different one.

Encoding is trickier still, because it does not show up in the structure but only in the content. If a file is written in a Western European encoding and read as UTF-8, accented characters turn into replacement symbols. Processing runs through, the columns are fine, and yet nobody can find those articles later. Faults like this move quietly into master data and resurface weeks later as an apparent search problem.

stock-2026-06-04.csv
# agreed: semicolon, UTF-8, date DD.MM.YYYY, article number four digits
article_no;description;stock;date
0041;Check valve 1/2 inch;12;04.06.2026
0042;Sealing ring 18 mm;340;04.06.2026

# actually delivered after an update of the source system:
article_no,description,stock,date
41,"Check valve 1/2 inch",12,2026-06-04
42,"Sealing ring 18 mm",340,2026-06-04

Three changes at once, all of them unannounced: a different delimiter, a lost leading zero in the article number, a different date format. An import that reads the delimiter from the file agreement and validates the header row rejects this file and reports the reason. An import that simply splits on commas because that used to work creates two new articles numbered 41 and 42.

Dates, numbers and leading zeros

Date and number values are the most common cause of faults nobody notices, because the result looks plausible. A date written as 06.05.2026 is the sixth of May; the same string read the American way is the fifth of June. Both are valid dates, neither produces an error message, and only one of them is correct. As long as the day figure is above twelve the mix-up is obvious; below that it is not.

  • Fix the date format and accept only that format on import. A deviating value is an error, not an invitation to guess.
  • Clarify the decimal separator. Comma and point mean different things depending on where the file comes from; a thousands separator in a quantity can turn 1,200 units into the value 1.2.
  • Protect leading zeros. Article numbers, postcodes and cost centres are text, not numbers. Read them as numbers and you lose the leading zero and hit the wrong row in the target system.
  • Check signs and units. A negative stock figure may be correct in business terms or may be a format fault; that decision belongs in the file agreement, not in the script.
  • Name the time zone and the cut-off. For overnight runs, the definition of the cut-off decides whether a posting still falls into the previous day or already into the current one.

These points sound pedantic but they are the heart of the matter. Almost every dispute about whose figures are right between two systems can be traced back to one of these five questions. Settling them once in writing saves the recurring hunt for the difference between two reports that were supposed to share the same basis.

The file agreement: putting the format in writing

The single most effective step costs no programming, just one page of text. A file agreement records what the file looks like, who delivers it, when it arrives and what happens if it does not. Both sides confirm it, and it is stored where somebody who does not work here yet can still find it. In data integration projects this document is regularly the point at which the real misunderstandings surface: two departments mean different things by the same column name.

Field list and order

Every column with its name, meaning, data type and whether it is mandatory. Plus the rule on whether the order is fixed or whether the header row governs. Both are acceptable, but only one of them may apply.

Delimiter and text qualifier

Which character separates, which one qualifies text, and how a qualifier inside a text is written. Plus whether a header row is present and whether trailing empty lines are permitted.

Encoding and line breaks

One encoding for all deliveries, one convention for line breaks. Plus a rule for characters that cannot be represented: replace them, reject the file or raise a report.

Date and number format

One date format, one decimal separator, a rule for thousands separators and for leading zeros. Complemented by the time zone and the definition of the cut-off for overnight runs.

Mandatory fields and value ranges

Which fields must not be empty, which values are permitted, which keys have to be known in the target system already. This turns a format check into a business plausibility check.

Delivery rhythm and expected volume

When the file arrives, what it is called, how many rows are normal, which deviation counts as suspicious. Plus the contact on the sending side, with a deputy and how to reach them.

The file name deserves more attention than it usually gets. A name with a fixed pattern and the date inside it, such as stock-2026-06-04.csv, makes duplicate deliveries, gaps and late catch-up deliveries visible at a glance. A name without a date that is overwritten on every delivery destroys exactly that information.

Duplicate deliveries and empty files

Two cases deserve separate treatment because they behave differently from classic format faults: the duplicate delivery and the empty file. Both are formally impeccable. Both pass every format check. And both do damage if processing does not recognise them as special cases.

A duplicate delivery happens faster than you might think: the overnight run was repeated by hand after an incident, a storage location was synchronised, a mail sat twice in the mailbox. If the import processes movement data, meaning receipts, issues or postings, the effect doubles. With stock data that contains the complete current state, a repeat is harmless by contrast. This distinction between movement and state belongs at the very start of any plan for process automation.

Repeating a run without double postings

The practical safeguard has two parts. First, a checksum over the file content: if it is already recorded in the log, the file is rejected and the repeat is merely noted. Second, a run identifier per business transaction that is written into the target system, so a row arrives only once even on a second attempt. With both safeguards in place a run may be repeated safely, and in an incident that is the property that matters.

The empty file is the second special case. It occurs when a report in the source system returns no hits, when an export aborts early or when a transfer is interrupted mid-write. Without a check, a stock import then sets all quantities to zero, because it reads the absence of rows as a statement about stock. That is why the file agreement needs a minimum volume and a rule for deviations: if the row count differs markedly from the usual volume, the file is not processed but reported. Completeness should also be recognisable, either through a trailer row carrying the row count or by having the sending side write the final file name only once the file is fully written.

Storage and logging: three folders and one entry per run

For storage, three folders have proven themselves. The inbox folder receives the file unchanged. After successful processing it moves to an archive folder, usually organised by year and month. If it is rejected, it goes to the error folder together with a small text file stating the reason. This split has an unspectacular but considerable advantage: the state of the import is visible without any tooling. Anyone who looks into the error folder sees immediately whether and where something is stuck.

Terminal
$ importrun --source stock --date 2026-06-04
06:05:02 File found: stock-2026-06-04.csv (18,412 bytes) 06:05:02 Checksum: 9f3c...a71b -- not processed before 06:05:03 Delimiter, encoding and header row: fine 06:05:03 Rows: 1,284 (previous day 1,291, deviation 0.5 percent) 06:05:06 Imported: 1,284 rows, 37 of them changed 06:05:06 File moved to archive/2026-06/ 06:05:06 Result: OK
$ importrun --source stock --date 2026-06-05
06:05:01 File found: stock-2026-06-05.csv (94 bytes) 06:05:01 Rows: 0 (previous day 1,284) -- below minimum volume 06:05:01 Processing aborted, target system unchanged 06:05:01 File moved to errors/2026-06-05/ 06:05:02 Alert sent to the distribution list 06:05:02 Result: REJECTED

The log is the second half. One entry per run with timestamp, file name, checksum, row count, result and the reason for a rejection is enough for daily operation. The IT baseline protection framework of the Bundesamt für Sicherheit in der Informationstechnik (BSI) treats logging as a module of its own and requires, among other things, that log data is collected for a defined purpose, protected against unauthorised access and deleted after defined retention periods (BSI). For a file import that mainly means: no personal content in the log, only figures about the run.

Whether the raw file itself has to be retained, and for how long, depends on whether it contains records relevant under tax or commercial law. For document and posting data that is regularly the case; for a plain stock list used only for display it is not necessarily so. This classification is a specialist question of the individual case and belongs in a discussion with your tax adviser; the technical side merely has to make sure the storage layout can reflect whatever was decided. Introducing an archive only after a tax audit means picking the least convenient moment.

Alerting: who learns what, and when

An alert nobody reads is not an alert. The most common design fault is the success message: if the import runs fine and still sends a mail every morning, the recipient will filter it away within two weeks, and the failure case along with it. Hence the simple rule that only deviations are reported. The successful run belongs in the log and in an overview you can call up when you need it.

  • A business alert to the process owner. It describes the consequence in the language of the company: today's stock list was not imported, the figures in the target system are yesterday's.
  • A technical alert to whoever maintains the routine. It contains file name, checksum, row count and the error message in plain text, so the cause can be traced without asking back.
  • A missing file as a case of its own. If nothing has arrived by the agreed time, the alert states that nothing has arrived. Without this check, a failed export on the other side stays invisible.
  • One alert per incident, not a repeating loop. If the same failure mode recurs across consecutive runs, a daily reminder is enough instead of a mail per attempt.
  • Name a deputy. A distribution list with two recipients and a named deputy survives holidays and illness; a personal mail address buried in a script does not.

The second point concerns the channel. Mail is convenient but awkward in an incident, particularly when mail delivery itself is affected. In companies with ongoing support, an overview page showing the state of recent runs therefore complements the alert. Whoever looks after IT operations sees at a glance which import last succeeded and when, without searching a mailbox.

File transfer or API?

The question is often treated as a matter of faith, but it is a trade-off. An API delivers data at shorter intervals, allows targeted queries and gives immediate feedback on success or failure. It does, however, assume that both systems are permanently reachable, that credentials are maintained and renewed, and that somebody responds when the provider changes versions. If one side goes down, you need retry logic and a queue, otherwise transactions are lost.

CriterionFile transferAPI
Timelinessdelivery rhythm, usually daily or hourlyclose to immediate, depending on the trigger
Traceabilitythe raw file remains as evidencerequires deliberate logging of request and response
Behaviour on failurethe file stays put, processing can be repeated laterthe call fails; without a queue data is lost
Dependence on the providerlow, almost any system can produce a file formathigh, version changes and access procedures need maintenance
Effort to get startedlow, often possible with the source system's own toolshigher, involving credentials, permissions and a test environment
Best suited tomaster data, stock levels, daily movements, reporting datatransactions needing instant feedback, such as an availability display

From this follows an unspectacular recommendation: a file import with solid error handling is preferable to a badly built API. It is easy to grasp, it leaves evidence behind, it can be repeated safely and it keeps working when a system is unreachable for two hours. Only when the business requirement calls for feedback within seconds does the API become the right choice. That decision is best made after a sober process analysis rather than on the question of what sounds more modern.

Rollout in four steps

Collect every running file handover: source, target, rhythm, storage location, responsible person, last known incident. In companies with a system landscape that grew over the years, this regularly turns up handovers the IT staff knew nothing about (field experience).

The parallel run in step three is readily skipped, and yet it is the step that builds confidence. Two to four weeks with both procedures side by side show whether the results agree, and they uncover exactly those special cases that appeared in no description: the month-end close, the catch-up delivery, the public holiday without a delivery.

Effort, cost and ongoing operation

For orientation: taking stock of existing handovers, with an assessment and a running order, starts at 1,900 euros net; implementing a single handover with check chain, storage, logging and alerting starts at 4,900 euros net. How far the figure moves upwards depends less on the technology than on the number of special cases in the department. An import whose rules fit on one page is done within days; one with twelve exceptions for individual customer groups is not. Current rates are listed on the pricing page.

The second cost block is operation, and it is underestimated more often than the build. A handover has to be watched, formats change with updates to the systems involved, contacts move on. Ongoing support starts at 190 euros net per month and covers monitoring the runs, handling alerts and adjusting to format changes. If you do not outsource that, at least name the task internally: an import without an owner is an import with an uncertain future.

This article is based on data from: Bitkom, the Bundesamt für Sicherheit in der Informationstechnik (BSI) and our own project experience from integration work in mid-size companies.

Related Articles

Automation & interfaces

Monitoring interfaces properly: beyond the server being up

Monitoring interfaces beyond availability: business-level checks, thresholds without false alarms, escalation to a named owner and logs with defined retention.

13 min read
Automation & interfaces

When an interface fails: spotting silent outages

Detecting, reporting and bridging silent interface failures: heartbeat, time windows, volume reconciliation, queueing and a defined restart procedure at work.

13 min read
Automation & interfaces

Webhooks instead of polling: when the switch pays off

Poll on a schedule or get notified? What each approach costs in daily operation, what a receiver must provide and when polling remains the safer choice.

13 min read