A document is handed over to the accounting system, the connection waits for an answer, and the answer does not arrive. After thirty seconds the attempt is cancelled, after five minutes the retry starts. What nobody sees at that moment: the first attempt did arrive and was posted, only the confirmation was lost on the way back. The result is two identical documents, two postings and one open item too many. Duplicate postings of this kind are rarely spectacular. They barely stand out in daily work because each document looks correct on its own. They surface at payment, at the VAT return or at the annual accounts, and by then the clean-up costs far more than the protection would have. This article explains why an unprotected retry produces exactly this outcome, how idempotency prevents it, where duplicates typically arise in an interface and in everyday routines, and how to find old cases in your records before someone else does.
Key takeaways
- A retry is not a mistake, it is correct behaviour — it only becomes dangerous when the receiving system cannot tell the second attempt from the first and posts the same document again.
- Idempotency means that the same operation may be transmitted any number of times and still takes effect exactly once; it rests on three parts: a unique operation key, a check before writing and a log of the keys already processed.
- The key must originate in the sending system, identify the business operation and stay unchanged across all attempts — a timestamp, a sequential transmission number or a random value per attempt are unsuitable.
- Duplicates in mid-size companies rarely come from technology alone: double clicks on batch buttons, staff working in parallel, files imported twice and the restart of an aborted nightly run are the most frequent triggers (project experience).
- Old cases are found by searching for equal amount, equal business partner and a close document date; confirmed duplicates are reversed and documented rather than deleted, and the assessment of each individual case belongs with your tax adviser.
How a retry turns into a second invoice
Two systems exchange data over a line that is not reliable. That is not a weakness of any particular software but a property of every connection running across a network. The sender transmits a document and waits for confirmation. If it arrives, the situation is clear. If it does not arrive, the sender knows only one thing: no answer received. It cannot tell whether the document was lost on the way, whether the target system received and discarded it, or whether it was posted there long ago and only the reply got stuck halfway.
This is the heart of the matter: no answer does not mean that nothing arrived. The sender still has to act, because staying silent would be worse. If it simply gave up whenever an answer failed to appear, documents would be lost and the two datasets would slowly drift apart. That is why every sensibly built handover repeats the attempt, usually with growing intervals: after one minute, after five minutes, after half an hour. So the retry is not the error. The error is that the other side does not recognise it as a repetition.
Without a distinguishing mark the target system simply sees a new document with identical content on the second attempt. It has no reason to be suspicious, because two invoices for the same amount to the same customer on the same day are perfectly possible in business terms. So it creates a second record, assigns a second internal number and reports success. The sender is satisfied, the target is satisfied, and accounting holds one amount twice. With a direct debit the amount is collected twice, with a stock movement the quantity leaves the warehouse twice, with a time entry the same hour appears twice on the project invoice.
Why a missing answer proves nothing
Idempotency: an awkward word for a simple promise
An operation is idempotent when it can be carried out any number of times and still takes effect exactly once. An everyday example: setting a delivery address to a new value can be sent ten times — afterwards the same address is stored, no matter how often it was transmitted. Creating a document or reducing a stock level by a quantity, by contrast, changes the outcome with every attempt. Such operations are not idempotent by nature and need protection added from the outside.
This protection consists of three parts that only work together. First, every operation needs a key that identifies it uniquely and stays the same across all retries. Second, the receiving system must check before every write whether this key has been seen before. Third, there has to be a log holding the keys already processed together with their result — because the check can only find what someone wrote down. If one of the three parts is missing, the protection is not half present, it is absent.
Operation key
An identifier assigned by the sending system that belongs to the business operation, not to the transmission attempt. It travels unchanged with every repetition. Only this allows the receiving side to recognise that it is looking at the same document for the second time.
Check before writing
Before a record is created, the target system looks into the log to see whether the key is known. If it is, nothing is posted; instead the stored result of the first attempt is returned. The sender receives a valid answer without a second posting being created.
Log of processed keys
A list of all processed keys with time of arrival, result and the internal document number assigned. It is the only place where it can later be proven what happened to an operation, and it is at the same time the tool for every subsequent investigation.
The effort involved is modest. In most projects it comes down to one additional column, one additional query and a table with a handful of fields (project experience). What becomes expensive is not the installation but retrofitting into a landscape that already holds duplicates and where nobody remembers which document came first. Anyone setting up a new automation should therefore plan for these three parts from the start, even if everything runs smoothly during testing.
The operation key: what it is made of and what it is not
A usable key has one decisive property: it does not change when the same operation is sent again. That rules out every candidate created at the moment of sending. A timestamp differs on the second attempt. So does a sequential transmission number. A random value generated by the interface per call turns every attempt into a new operation and is therefore the exact opposite of what is needed.
The key has to come from the business operation itself, which means from the sending system where the document was created. A combination of document type, document number and a short code for the source system has proven useful, for example in the form INV-2026-4417 for an outgoing invoice. This identifier is already unique in the source system, it is readable for people, and it also appears on the printed document during a later clarification. Alternatively a technical identifier can be used, generated when the operation is created and stored permanently on the record.
- The key is assigned when the operation is created, not when it is sent, and stays stored on the record permanently
- It is unique across all document types so that invoice 4417 and delivery note 4417 cannot collide
- It contains no data that can change in business terms, such as customer number, amount or processing status
- It travels unchanged with every retry, including after a restart of the sending service
- It is readable and pronounceable so that it remains usable during a clarification on the phone
- It is stored on both sides so that the link can still be traced years later
A frequent mistake in practice is trying to do without a key and comparing content instead: same customer, same amount, same date, so probably a duplicate. This works surprisingly well for invoices with unusual amounts and badly for everything else. Two maintenance flat rates for the same amount on the same day are an entirely normal transaction. Content comparison as the only safeguard therefore leads either to missed duplicates or to rejected genuine documents. As an additional search aid for old cases it is useful; as protection during live operation it is not.
The check before writing
The second ingredient is a query that sits in front of every write access. It answers exactly one question: has this key been processed before? If the answer is no, the key is reserved first, then the document is posted and finally the result is recorded against the key. If the answer is yes, nothing is posted. Instead the sender receives the stored result of the first attempt, meaning the same internal document number it would have received the first time.
This last point is often overlooked. Silently discarding the second attempt is not enough, because the sender is then left in the dark and will try again. The answer to a recognised retry is not an error message but a calm confirmation stating that the operation already exists. Only then is the loop closed and the document known as completed on both sides.
1 document arrives, operation key: INV-2026-4417
2 look up in the log: is the key known?
3 no -> reserve the key (unique index in the database)
4 -> post the document
5 -> record the result against the key: document 88231
6 -> answer to the sender: created, document 88231
7 yes -> do not post again
8 -> read the stored result: document 88231
9 -> answer to the sender: already present, document 88231Between step two and step four lies a short window in which a second attempt can arrive. With fast repetitions or services working in parallel this happens more often than expected. Uniqueness must therefore not depend on application logic alone but also belongs in the database: a unique index on the key column. If the second attempt starts anyway, it fails at the reservation, and the application can treat this case cleanly as a repetition. Without that safeguard a narrow gap remains, and it opens precisely under load.
The log of processed keys
The log is the least conspicuous of the three ingredients and the most important during clean-up. For each key it records when the operation arrived, how it was decided and which internal number was created. A checksum over the transmitted content and the name of the sending system are useful additions. The checksum answers a question that otherwise stays open: was the same key sent with different content? That is not a retry but an indication of an amended invoice or of a fault in the way keys are assigned, and neither should be rejected silently.
There is no universally valid figure for the retention period. It has to cover at least the window in which retries are possible, plus the time during which someone may want to clarify an operation. In practice a minimum of ninety days has proven workable, and considerably more on document-bearing routes (project experience). For logging itself the German Bundesamt für Sicherheit in der Informationstechnik (BSI) offers orientation in its baseline protection modules: scope, storage location and retention should be defined in writing and reviewed regularly (BSI). Where logs relate to tax-relevant operations, commercial and tax retention periods apply in addition, and their application in each case should be checked with a professional.
The log is also evidence
Where duplicates typically arise in mid-size companies
The technical retry is only one of several routes. In practice a substantial share of duplicates arises where people and systems meet, and these places are the same in almost every company. To assess your own situation, walk through the following table once with accounting and once with warehouse management. Experience shows that such a conversation surfaces more cases than any report on screen (project experience).
What stands out is that most triggers involve an interruption: a screen that hangs, an aborted nightly run, a colleague working on the same item in parallel. Duplicates rarely arise during calm normal operation; they arise almost always when something does not go as planned and somebody steps in.
| Place of origin | How the duplicate arises | What helps |
|---|---|---|
| Batch button in inventory management | The screen responds slowly, the user clicks transfer a second time | Disable the button after the first click, assign an operation key per document |
| Retry of the interface | The answer is lost, the document is sent again and posted again | Check before writing and a log of processed keys |
| File import from an upstream system | The same file is read in a second time after a query | Log file name and checksum, reject files already processed |
| Aborted nightly run | The run is restarted and begins again at the top of the list | Store a restart point, check a key for each record |
| Parallel processing by staff | Two people post the same receipt because no lock is held on the item | Lock the item when it is opened, make the processing status visible |
| Manual entry after a fault | The operation is entered by hand and later runs through automatically as well | Mark manual entries, review the queue before restarting |
The last row is particularly unpleasant. After a fault people readily enter by hand what was left undone, and when the connection returns it works through its queue with exactly those operations. Anyone running a data handover between systems should therefore define in writing who enters items manually during a fault and what happens to the queue before the restart. This agreement costs nothing and prevents one of the most frequent causes of duplicates.
Finding old cases before your tax adviser does
Retrofitting the protection solves the problem for the future and not for the past. The duplicates already created remain in your records, and they surface at the least convenient moment: during the reconciliation of open items, during an audit or when a customer asks why an invoice arrived twice. A targeted search through existing records is therefore the second half of the task, and it can be structured in a few steps.
The search approach is the very one that would be unsuitable as live protection: content comparison. As an investigative tool for a closed period it works well, because every hit is assessed by hand anyway. The only important thing is to keep the result list generous and sort it afterwards, rather than setting the criteria so narrowly that an empty list provides the relief.
Step 1: define period and document types
A sensible period is the time during which the affected connection ran without protection, but at least the current and the previous financial year. It is also defined which document types are examined: outgoing invoices, incoming invoices, stock movements, payments. Each type gets its own list because the assessment differs.
Step 2: pull candidates with wide criteria
The search looks for document pairs with the same business partner, the same gross amount and a document date no more than a few days apart. The range is deliberately wide so that duplicates that ran through with a delay also appear. The result is a candidate list, not yet a finding.
Step 3: enrich and sort the candidates
For each pair the line items, the document number of the other side and the time of creation are added. Pairs with identical line items and a creation gap of a few minutes go to the top, because they are the most likely hits. Pairs with differing line items move to the bottom and are usually genuine transactions.
Step 4: assess in business terms, not technical ones
The assessment is done by accounting, not by technology. For each pair it is noted whether it is a duplicate and, if so, which document is the valid one. This decision is documented because it later forms the basis for reversal and correction and must be explainable to third parties if questioned.
Step 5: discuss the result with your tax adviser
Before anything is corrected, the list belongs on the desk of your tax adviser. They decide on the route of reversal, the timing and whether a return already submitted is affected. This order saves rework, because a correction carried out alone may have to be corrected again.
A sensible period is the time during which the affected connection ran without protection, but at least the current and the previous financial year. It is also defined which document types are examined: outgoing invoices, incoming invoices, stock movements, payments. Each type gets its own list because the assessment differs.
The search looks for document pairs with the same business partner, the same gross amount and a document date no more than a few days apart. The range is deliberately wide so that duplicates that ran through with a delay also appear. The result is a candidate list, not yet a finding.
For each pair the line items, the document number of the other side and the time of creation are added. Pairs with identical line items and a creation gap of a few minutes go to the top, because they are the most likely hits. Pairs with differing line items move to the bottom and are usually genuine transactions.
The assessment is done by accounting, not by technology. For each pair it is noted whether it is a duplicate and, if so, which document is the valid one. This decision is documented because it later forms the basis for reversal and correction and must be explainable to third parties if questioned.
Before anything is corrected, the list belongs on the desk of your tax adviser. They decide on the route of reversal, the timing and whether a return already submitted is affected. This order saves rework, because a correction carried out alone may have to be corrected again.
A handover is only finished once it is described what happens on the second attempt.
What happens to a duplicate once it is found
The obvious reaction to a duplicated document is to delete it. That is precisely what should not happen. Documents that have been posted are reversed, not removed, so that the sequence of postings remains traceable. A deleted document leaves a gap in the numbering that has to be explained during any later review, and years later nobody remembers the explanation.
The second point concerns the order of work. Before a duplicate is corrected it should be clear whether it had consequences: a payment triggered, a reminder sent, a stock movement, a report transmitted to a third party. These consequences are listed first and then handled together. Reversing only the document and leaving the triggered payment in place merely moves the problem into payment reconciliation.
- Reversal instead of deletion, with a reference to the valid document in the posting text
- List the consequences before correcting: payments, reminders, stock movements, reports to third parties
- Inform affected business partners actively when a document has gone out of the house
- Record the cause in the same case file so that the retrofit targets the right place
- Agree tax-relevant cases with your tax adviser before correcting
- Keep the correction list so that the clean-up remains provable later on
The legal and fiscal assessment of a specific case is not a task for technology. This section describes a factual way of proceeding and does not replace legal or tax advice; the assessment of the individual case belongs in qualified hands, particularly where submitted returns or already closed periods are affected.
Questions to ask about an existing connection
Whether an existing handover is protected can be clarified without access to the source code. A few questions to whoever looks after the connection are enough — internally or at your service provider. What matters is less the answer itself than its form: anyone who can name the operation key and show the log probably has the protection in place. Anyone who replies that such a thing has never happened probably does not.
- Which field carries the operation key, and who assigns it — the sending or the receiving system?
- Does this key stay unchanged during a retry, including after a restart of the service?
- Where is the log of processed keys kept, and how long are the entries retained?
- What does the target system answer for a key it already knows: an error message or the first result?
- Is there a unique index in the database, or does the check rely on application logic alone?
- When was the same document last sent twice on purpose, and what was the outcome?
The last question is the most revealing, because it cannot be answered with a description. A deliberately repeated test document in a test environment takes a few minutes and produces an unambiguous result: afterwards the document is either present once in the target system or twice. Everything else is interpretation. Anyone running connections regularly adds this test to the recurring tasks of day-to-day IT operations, alongside checking the queue and the size of the log.
A starting point without a project
Related Articles
Duplicate data entry: three routes out of retyping
Typing the same data again and again: why duplicate entry arises, what it costs and the three routes out: a leading system, an interface, or file transfer.
Invoice checks automated: order, goods receipt, invoice
Matching order, goods receipt and invoice by machine: which fields are compared, where the tolerance band sits, who owns the exception and what stays manual.
Replacing the grown spreadsheet: when and how it pays off
When a spreadsheet becomes a shadow core system: five weak points, three decision questions and a transition path with a parallel run instead of a standstill.