Skip to content
Data & documents

Migrating Data from a Legacy System Without Surprises

Inventory, field mapping, trial runs and totals checks: how to move data out of a legacy system and how to evidence that the transfer really was complete.

14 min read DatenmigrationAltsystemeSystemwechselDatenqualitätVerfahrensdokumentation

A system change rarely fails on the new software. It fails on what is supposed to come across from the old one: twelve years of order history, customer master data in three spellings, article numbers with leading zeros and a comments field that absorbed everything for which no dedicated field existed. Treat data migration as the last item on the project list and you will discover on cutover weekend that it was the single largest risk. This article describes the approach that works in mid-size companies: take stock before touching technology, select deliberately rather than moving every record, document the field mapping, run at least two trial migrations, reconcile the totals verifiably and plan a cutover date with a proper follow-up phase. At the end stands the question an auditor may ask: what evidence do you have that the transfer was complete? How a legacy system is replaced in an orderly way depends on that answer.

Key takeaways

  • Data migration is the part of a system change with the highest risk and the shortest lead time: it is usually planned last, even though it determines whether the cutover date holds.
  • Before any technical work comes the inventory: which data sets exist at all, who owns them in business terms, how many records they contain and which of them are still moving in day-to-day work.
  • Leaving data behind is not a weakness — dormant records, duplicates and test data make the migration more expensive and burden the new system; historic volumes usually belong in a read-only archive rather than in the productive data set.
  • The field mapping belongs in writing, field by field, including conversion rules and the handling of fields with no counterpart; it also forms the basis of the later process documentation.
  • Completeness is evidenced on three levels: number of records, business totals such as open items or stock values, and samples of individual cases — every unexplained deviation pushes the cutover date back.

Why data migration is the real risk

In most replacement projects, choosing the new system is the part that gets discussed longest. Feature lists are compared, demonstrations watched, quotations reviewed. Data migration appears in the quotation phase as a single line: data migration, time and materials. That very line later decides whether the planned date holds. The reason is simple: the new system can be demonstrated, your own data set cannot — only the person who opens it and looks at it row by row really knows what is in there.

Data sets that have grown over years are rarely as tidy as everyone assumes. Workarounds accumulate: a customer is created twice because the search did not find them. A special condition sits in the comments field because the system has no field for it. An article group is repurposed to make a report possible. None of this was done wrongly — it was the path the old system left open. During migration, however, this organically grown data meets a system with stricter rules, and there every workaround shows up individually.

There is a simple reason such changes are becoming more frequent: specialised applications are increasingly used for individual tasks, while older all-in-one solutions reach the end of their support. The European Commission has set a target for 2030 that 75 percent (European Commission) of companies in the EU use cloud services, data analytics or artificial intelligence. In Germany more than 99 percent (Statistisches Bundesamt) of companies are small and medium-sized enterprises — they carry out such changes without a dedicated project department, usually alongside daily business. That makes an approach without a project office all the more important.

Three terms that get mixed up

Data migration means the one-off transport of a data set from an old system into a new one at a cutover date. Data cleansing means tidying that data set up: merging duplicates, standardising spellings, correcting invalid values. Data integration means the ongoing synchronisation between systems that keep running in parallel. The three tasks overlap in a project but should be planned and signed off separately.

Taking stock: what is actually in the legacy system?

The first step is unspectacular and gets skipped regularly all the same: a list of every data set that matters for operations. Mattering does not mean sitting in the main database. In practice an order system comes with side records that nobody calls a system: a spreadsheet of special prices, a folder of signed contracts, a mailbox with supplier confirmations, a hand-kept list of contact people. Overlook them and the gap appears right there after cutover, because the new system takes that data for granted.

For each data set found, a handful of details are recorded. They sound trivial but save follow-up questions later and are what makes the effort estimable in the first place. In our experience this inventory takes one to three working days in a company with one main system and a few side records — considerably less than clarifying the same questions under time pressure shortly before cutover (project experience).

  • Name and location: in which system or which filing place does the data set actually sit?
  • Business ownership: who decides what is correct when two values contradict each other?
  • Volume: how many records are there, and how many of them have moved in the last two years?
  • Dependencies: which other data sets reference it, such as orders referencing customers or lines referencing articles?
  • Export route: can the data set be output in an orderly way, and in which format and character set?
  • Obligations: is the data set subject to retention rules, or does it contain personal data?

The last question tends to be deferred and belongs at the front. Personal data on staff, applicants or customers may not simply be copied into a new system just because it is technically available; purpose limitation and deletion periods continue to apply. Likewise, records under tax or commercial retention obligations deserve separate treatment. Both should be examined professionally in the individual case — this article is not legal advice. Anyone carrying out the inventory as part of a process analysis usually gets the answers as a by-product.

Not everything comes across — and that is no loss

The common reflex is: let us take everything, just to be safe. It is understandable and expensive. Every additional record has to be mapped, converted, checked and, in case of error, clarified individually. A data set full of dormant records generates more work during checking than the active part, because it is precisely the old records that are the most incomplete: missing VAT identification numbers, addresses without a country, articles without a unit of measure. These records surface as errors in the trial run and consume time that the active data then lacks.

On top of that comes the permanent effect: whatever is migrated appears afterwards in every search, every selection list and every report of the new system. A sales team scrolling past three thousand inactive entries during a customer search loses time on every working day — and acceptance of the new system drops during the first weeks, when it matters most. Selection is therefore not a cost-cutting measure but a quality decision. It should be made by the business side and justified in writing, so that it remains traceable later why a data set is not in the new system.

Type of dataUsual decisionReason
Open orders and purchase ordersMigrateThey continue to be processed in the new system after cutover
Stock levels at the cutover dateMigrateBasis for planning and inventory value from day one
Master data with activity in 24 monthsMigrateActive data, needed in daily business
Running contracts, prices, conditionsMigrateWithout them, incorrect documents are produced from cutover onwards
Documents under retention obligationsMigrate or archive in an audit-proof wayAccess has to remain possible for the entire retention period
Master data dormant for yearsLeave behind, re-enter individually if neededCreates checking effort and clutters every later search
Duplicates and test recordsClean up before migratingOtherwise they come across and become entrenched in the new system
Order history of earlier yearsInto a read-only archive rather than the productive dataAbility to answer questions remains without burdening the new system

Field mapping: from column to column

The field mapping is the heart of the migration and at the same time the part that most often stays verbal. For every field of the old system it answers three questions: which field of the new system does the content go into? Does it have to be converted or reshaped on the way? And what happens if the field is empty or invalid? As long as those answers live only in one person's head, the migration is neither verifiable nor repeatable — and repeated it certainly will be, because no trial run passes on the first attempt.

In practice the mapping is kept as a table with one row per field. The conversion rules matter most, because that is where the actual work sits: removing or preserving leading zeros, splitting or joining first and last names, standardising date formats, amounts with a point or a comma, converting units of measure, truncating text lengths. Each of those rules is small on its own. Together they decide whether the data is analysable in the new system or whether the same workarounds reappear there within two years.

Extract from a field mapping (customer master data, example)
Legacy system              New system             Rule
-------------------------------------------------------------------------------
KUNDENNR (text, 8)         customer.number        strip leading zeros
NAME1 + NAME2              customer.name          join, separated by a space
STRASSE                    address.street         split house number into its own field
PLZORT (text, 40)          address.zip / .city    split at the first space
USTID                      customer.vat_id        remove spaces, upper case
ZAHLZIEL (number, days)    customer.payment_term  unchanged
SPERRKZ (X or empty)       customer.blocked       X becomes true, empty becomes false
BEMERKUNG (free text)      customer.note          migrate, do not evaluate automatically
RABATTGRUPPE               —                      no counterpart, see open items list

A second list belongs with the mapping: the open items list. It holds every field for which no decision has been taken yet, each with a responsible person and a deadline. A migration whose open items list still has entries on the day of the trial run is not ready for a cutover date. Which system leads for which field afterwards is a separate question, by the way: it belongs to data integration and should be settled before the first record is copied.

Fields with no counterpart: four possible routes

Every project has fields for which the new system provides no place. That is the moment when migrations derail — not because the decision is difficult, but because it gets postponed. Four routes are available, and each is right in certain cases. What matters is only that one of them is deliberately chosen and recorded for every affected field.

Move into an existing field

Often the new system has a field with a similar meaning, such as a category instead of a numeric code. A translation table is then created: old value on the left, new value on the right. This is the cleanest route because the information remains analysable.

Create an additional field

If the information is genuinely needed and has no counterpart, a dedicated field is set up in the new system. That costs effort and should be justified, otherwise the same sprawl grows there over time as in the old system.

Preserve it as a note

For content that is only read occasionally and never evaluated, a note on the record is enough. The content is not lost but is deliberately not analysable — everyone involved should be clear about that beforehand.

Deliberately leave it out

Some fields were created years ago and have not been maintained since. They are not migrated, and that decision is recorded with a date and a responsible person. The old data set stays readable until the end of its retention period.

Free-text fields deserve particular attention. In systems that have grown over years they often hold information that really belongs in structured form: payment arrangements, delivery instructions, access details for customer portals, names of contact people. Simply carrying them across as a note is the fastest route, but it defers the problem. An intermediate step helps: the free texts are read, recurring patterns identified and two or three structured fields derived from them. The rest travels along as a note. That way the valuable part of the free text is rescued without every sentence having to be assessed individually.

Historic data: migrate, archive or leave in place

With history, two legitimate interests collide. In business terms you want to be able to answer questions: what did this customer buy four years ago, at what price, with which serial number? In technical terms you do not want to fill the new system with records that are only ever read. The answer is rarely all or nothing but a split: a limited period moves into the productive data, the older part into a read-only archive.

Such an archive can be very plain. Often an exported, searchable data set with documents as files and an overview table containing document number, date, customer, amount and file name is enough. What counts is not the technology but that access is governed: who may search, where does the data sit, how is it backed up, and until when does it have to remain available? If documents exist on paper or scattered around anyway, it pays to combine this step with digitising the documents rather than handling them twice.

The retention period applies to the data, not to the old program

Retention obligations relate to records and their readability, not to keeping a particular program running. The periods differ by type of record and were adjusted for certain document types recently — which period applies to which data set, and in what form data has to remain machine-readable, belongs in a tax review of the individual case. In practice this means: before a legacy system is switched off, it must be clear how the data under retention obligations stays readable and findable afterwards.

The trial run: at least twice, on real data

A trial run on invented test data proves that the procedure works technically. It does not prove that it works with your own data — and that is the point. The trial run therefore has to use a genuine copy of the legacy data, in an environment separated from live operations. Only there do the unforeseen cases appear: the customer with an eight-line address, the article with a negative quantity, the date from 1899 that some program set as a placeholder decades ago.

A single trial run is not enough. The first one uncovers the coarse errors; after the corrections, the second shows whether those corrections work and have not created new errors. Only once a run completes without unexplained deviation and the business checks have passed is the cutover date fixed as binding. Anyone promising the date before the second clean run ends up negotiating about errors instead of about timing (project experience).

A complete copy of the legacy data with a recorded timestamp. All following figures refer to exactly this copy, otherwise no reconciliation is possible because the live data keeps moving.

A frequent mistake is separating the technical check from the business check. The technical side verifies that all rows arrived. Whether the value in the field still means what it meant before can only be judged by the business department. An example: if the discount field in the legacy system was kept as a percentage and in the new system as a factor, all rows arrive, all row counts match — and every price is nevertheless wrong. Only recalculating real cases finds that.

Reconciling the totals: how completeness is evidenced

Completeness cannot be asserted, it has to be shown. What works is evidence on three levels, collected both before and after the migration. First the count: how many records does the data set hold in the source, how many in the target? Second the business total: what is the value of the open items, what is the value of the stock, what is the total of open orders? Third the sample: twenty randomly drawn cases are compared field by field.

Differences are not forbidden — unexplained differences are. If thirty records are missing because they were merged as duplicates, that is fine, provided the number thirty follows from the cleansing list and the sum of values stays the same. If thirty records are missing and nobody can say which, the migration is not signed off. That distinction is the whole trick: what counts is not a zero in the difference column but the explainability of every deviation.

Terminal
$ reconcile --dataset customers --source legacy --target new
Rows source: 4,812 Rows target: 4,788 Difference: 24 Merged as duplicates: 24 unexplained: 0
$ reconcile --total open-items --cutoff 2026-06-30
Total source: 412,806.44 Total target: 412,806.44 Difference: 0.00
$ reconcile --sample orders --count 20
20 cases checked, 20 matching, 0 open deviations
  1. Row count per data set in source and target, each with the time of measurement.
  2. Total of open items from receivables and payables at the cutover date.
  3. Total and count of stock levels, split by location if several are kept.
  4. Number of open orders and purchase orders together with their total value.
  5. Sample of at least twenty cases across all types of data, compared field by field.
  6. List of all data deliberately not migrated, with reason, date and responsible person.

Cutover, parallel operation and follow-up

The cutover date is the point from which work happens in the new system. It should fall into a quiet phase and ideally coincide with a period close, so that the totals are comparable without intermediate calculations. A month change is enough in many companies; in retail or logistics the seasonal situation matters more than the calendar. Before cutover the legacy data is frozen: no further postings in the old system, changes from that moment on happen only in the new one.

A frequent wish is parallel operation: both systems run alongside each other for a few weeks so that you could fall back if needed. It sounds like safety and is demanding in practice, because every posting has to be entered twice. If that is not maintained with discipline, two versions of the truth emerge and nobody knows which one applies. A middle form carries better: the legacy system stays available for reading, while the new system alone is authoritative. Anyone who genuinely needs a fallback option should time-limit it and define in advance the condition under which the switch back happens.

A parallel run without a fixed end is not a safety net but a double workload with two data sets drifting apart. Safety comes from the verified trial run beforehand, not from the open back door afterwards.

Project experience
  • Plan for stragglers: documents arriving after cutover but carrying an earlier date need a defined route.
  • Corrections to old cases: who may look things up in the read-only archive, and how is a subsequent credit note represented in the new system?
  • Number ranges: sequential document and customer numbers must not overlap; the starting value in the new system is set before cutover.
  • Interfaces: every connection to other systems has to point at the new system on the cutover date, including rarely used ones.
  • Availability: the first days need named contact people for questions, otherwise workarounds appear and stay.

The follow-up phase does not end with the first working day without incidents. A review after four to six weeks is worthwhile: do the reports add up? Are cases appearing that nobody can place? Are there fields that have stayed empty since cutover because they no longer occur in the workflow? That review costs a few hours and prevents an error from surfacing only at the year-end close, when nobody remembers where it came from.

The evidence: what has to be documented at the end

At the end of a migration stands a migration record. It is not a formality but the answer to a question that may be asked in two or five years — by an auditor, by an accountant, by a new colleague wanting to know why a data set ends where it does. The record summarises which data sets were migrated when, from which source into which target, which figures were compared, which deviations occurred and how they are explained.

It includes the field mapping as it stood at the actual migration, the list of data deliberately not migrated together with the reasons, and the location and access rules of the read-only archive. These documents also form building blocks of the process documentation that is required for tax-relevant workflows anyway. Write them along during the project and you have them afterwards; try to produce them retrospectively and you are reconstructing decisions from memory. Orderly process documentation turns a project result into sound evidence.

The sentence everything can be measured against

A data migration is finished when an uninvolved person can follow, from the documents alone, which data was migrated, which was not, why the figures deviate and where the data left behind is still readable today. As long as that sentence does not hold, the project is technically done and unfinished in business terms.

A practical way to start

Start by counting rows, before any conversation about technology. For each data set, count the records and the records with activity in the last 24 months. Within a few hours that pair of numbers shows how large the migration really is — and it is at the same time the first measurement point for the later reconciliation. The German information security authority BSI recommends verifying backups regularly through restore tests (BSI); before a migration is exactly the right moment for such a test.
This article is based on data from: European Commission (Digital Decade targets for 2030), Statistisches Bundesamt (business structure in Germany), the German Federal Office for Information Security (IT-Grundschutz, backup and logging) and our own project experience.

Related Articles

Practice & rollout

Replacing the grown spreadsheet: when and how it pays off

When a spreadsheet becomes a shadow core system: five weak points, three decision questions and a transition path with a parallel run instead of a standstill.

14 min read
Law, security & funding

Writing process documentation for tax audits

Process documentation under the German GoBD rules: the four required parts, how detailed it must be, how to keep it current and what its absence can mean.

14 min read
Data & documents

Retention periods digitally: schedules, holds, deletion runs

Retention duties and deletion duties only appear to conflict. How to build a filing concept with periods per record type, legal holds and documented runs.

14 min read