Skip to content
Data & documents

Getting master data in order: customers, articles, suppliers

Customers, articles, suppliers: which system leads, which fields are mandatory and how to clean up master data without halting your daily operations.

14 min read StammdatenDatenqualitätFührendes SystemDublettenDatenpflege

Master data is the information that stays the same across many transactions: customers, articles, suppliers, along with numbers, units, prices and terms. It rarely takes centre stage, because nobody sells a customer record in the morning. Yet almost everything depends on it. Every report groups by it, every interface recognises records through it, every invoice and every purchase order draws on it. Where the same company is held under three spellings, where an article number was reissued after a deletion, or where nobody can say which system is right about an address, errors appear in the documents while their cause sits in the master record. This article describes how to sort that out: which system leads for each data type and each field, which fields really need to be mandatory, how spellings get standardised and who maintains which field. It then sets out a clean-up approach that works in waves and does not halt daily operations. If you want to measure your own data first, the starting point is the process analysis.

Key takeaways

  • Master data is the information that persists across many transactions; an error there does not occur once but repeats in every document, every report and every transfer created afterwards.
  • Reports fail on inconsistent keys, because a customer held under three numbers appears as three mid-sized accounts instead of one large one; interfaces fail because without a unique key they create new records instead of matching existing ones.
  • A system leads per field, not per data type as a whole: customer address in the merchandise management system, payment terms in accounting, bank details in purchasing under dual control, recorded in a table of field, system and reason.
  • Mandatory fields and spelling rules belong in the entry screen, not in a document in a folder; too many mandatory fields produce placeholder values, while fixed value lists instead of free text prevent most later reporting errors.
  • The clean-up runs in five steps and in waves: set the rules first, then secure new records, then measure the existing data, then clean in order of impact and finally take the same measurements again.

What master data is and why it belongs to nobody

A customer has a number, a name, an address, payment terms and a tax classification. An article has a number, a description, a unit of measure, a price and a product group. A supplier has a number, an address, terms and bank details. This information persists across many transactions and changes rarely. Set against it is transaction data: quotations, orders, delivery notes, invoices, postings. Transaction data is created afresh every day and refers back to the master record. That is where the leverage comes from: a wrong value in the master record does not take effect once, it takes effect in every document created afterwards and in every report built on those documents.

In daily work the master record has no advocate. An order has a delivery date, an invoice has a due date, a master record has neither. It is created on the side, usually under time pressure, because an order has to be entered right now. Whoever creates it fills in what is needed for their own purpose and leaves the rest empty. That is not carelessness but understandable prioritisation. Master data belongs to no department, so each one maintains it by its own standards. Only in a report, or when two systems are connected, does it become apparent that those standards do not fit together, and by then the data is several years old.

The vast majority of companies in Germany are small and medium-sized enterprises (Federal Statistical Office of Germany). They rarely have a position that owns data quality as a task of its own, and creating one would not be economic. The way out is not a new department but a handful of decisions that hold up in daily work: clear responsibility per field, a manageable number of mandatory fields, and rules that sit in the entry screen rather than in a document nobody opens. The following sections describe exactly those decisions and the order in which they make sense.

Three terms kept apart

Master data describes objects that persist across many transactions: customers, articles, suppliers, general ledger accounts. Transaction data describes individual events: order, delivery, invoice, payment. Reference data means value lists to choose from: units of measure, country codes, payment methods, tax codes, product groups. A substantial share of all reporting errors originates in reference data, because free text can be typed where a fixed list should stand.

Why reports and interfaces fail on master data

A report is essentially a grouping. Revenue per customer, contribution per product group, purchase volume per supplier: each of these figures is produced by aggregating documents over a shared attribute. If that attribute is inconsistent, the group falls apart. A customer held under three numbers appears as three mid-sized accounts instead of one large one and drops out of every ranking based on size. A product group left empty on part of the article base creates a catch-all position of unassigned items that may well be the largest block in the report. The report is not miscalculated; it simply answers a different question than the one that was asked.

With interfaces the same shortcoming hits harder, because no human sits in between. A transfer needs a key by which both sides recognise the same record: a customer number, an article number, a supplier number. If that key is missing or is not unique, the comparison falls back on name and address, and then the spelling decides whether a record is found or created anew. Duplicate records appear in the target system that nobody intended and that later have to be merged by hand. An interface does not check whether the data is right, it merely distributes it faster. Which connections are technically feasible is described under interfaces; the sequence still holds: master data first, then the connection.

The error shows in the document, the cause sits in the master record

When an invoice goes out with the wrong address, it gets corrected in the invoice, because that is where the pressure is. The master record stays as it was, and the next document repeats the error. This loop is a common reason why the same correction comes up again and again. A simple rule helps: whoever corrects an error in a document checks the underlying master record in the same step and changes it where it is led.

The most expensive effect is not the individual error but the loss of trust. When two reports for the same period show different figures, the cause is rarely investigated; more often the report is set aside and decisions go back to experience. Data the company has already paid to collect then goes unused. Anyone building a reliable reporting function should therefore start with the master data rather than with the reporting layer; which metrics carry weight in a mid-size company and where they come from is described under metrics and reporting.

One leading system per data type and per field

The first decision costs no development, only a choice: which system holds the binding version for which field? All other systems may display and use the value but must not change it on their own. That sounds like a formality yet changes a lot in daily work. Anyone wanting to correct an address then knows without asking where to do it, and nobody has to debate which of two differing versions applies. Without that clarity you get the familiar situation in which a change is made wherever the error was noticed and stays untouched everywhere else.

What matters is the level of the decision. A system leads per field, not per data type as a whole. Otherwise a debate starts about which system is the most important in the company, and that debate can neither be won nor put to use. Per field the question can be settled on the merits: a value belongs where it arises professionally and where it is checked in case of doubt. The address arises in sales and is led in the merchandise management system, payment terms are owned by accounting, a supplier bank account by purchasing and under additional control. The table below covers the cases that regularly turn out to be contentious.

Data type or fieldTypically leadingCommon disputeWorkable rule
Customer master, addressMerchandise managementWho may change the address?Change only there; documents keep their historical address
Payment terms, conditionsAccounting systemSales promises, accounting carries the riskCommitments only from a stored value list
Article master, commercialMerchandise managementWho maintains price and unit?Commercial fields centrally, texts and images in the department
Prices and discountsMerchandise managementSpecial prices in side listsModel pricing in the system, no lists alongside
Supplier master, bank detailsPurchasingWho may change bank details?Changes only under dual control with evidence
Personal reference in operationsTime recordingPersonnel number or initials?One identifier, the same in every system

Record the outcome in a table with three columns: field, leading system, reason. That table is not an administrative document but the basis for every later technical decision, because a transfer whose direction has not been settled can be modelled neither as an interface nor as a file handover. Where the system allows it, downstream fields should be write-protected; where that is not possible, a visible marker in the screen helps. The wider frame in which such decisions come together is described under data integration.

Mandatory fields: as few as possible, as many as needed

Mandatory fields are the most effective tool against gaps in the data and at the same time the one most often overdone. Declare twenty fields mandatory and you do not get complete records, you get placeholder values: a full stop in the name field, a zero in the phone field, a date from the past. Such values are worse than empty fields, because later on they can no longer be told apart from genuine entries. Choosing the mandatory fields is therefore a business decision rather than a technical one: mandatory is whatever makes a later step impossible or wrong if it is missing.

Besides mandatory and optional there is a third, often overlooked category: conditionally mandatory. A VAT identification number is needed for a business customer in another European country and not for a private customer at home. A named contact makes sense for a wholesale account and is superfluous for a walk-in customer. If the entry screen can reflect that distinction, the mandatory fields stay few and still work. If it cannot, a visible status on the record is a better answer than a blanket obligation for everyone.

Extract from a field catalogue, data type customer
Field              | Mandatory  | Leading      | Rule
-------------------|------------|--------------|------------------------------------
Customer number    | yes        | Merch. mgmt  | sequential, never reissued after deletion
Company name       | yes        | Merch. mgmt  | spelling as in the commercial register
Street, number     | yes        | Merch. mgmt  | house number in the same field, no space before suffix
Post code, town    | yes        | Merch. mgmt  | town without district in brackets
Country            | yes        | Merch. mgmt  | from a fixed value list, no free text
VAT identification | conditional| Accounting   | mandatory for business customers abroad in the EU
Payment terms      | yes        | Accounting   | in days, only from the stored value list
Named contact      | no         | Sales        | free text, no collective addresses in the name field

# mandatory means: without this value no record is created
# conditional means: mandatory only in a described case, otherwise deliberately empty
# the catalogue belongs in the entry screen, not in a folder

For electronic invoicing there is a European format that prescribes certain details (European Commission). Fields required there consequently belong in the mandatory part of the customer master, because otherwise they are missing at the very moment an invoice is issued and then have to be obtained under time pressure. The same applies to details your warehouse or production needs: a unit of measure left empty in the article master leads to a query at picking at the latest. Check every field against the following questions before declaring it mandatory.

  • Without this value, does a later step become impossible, or merely less convenient?
  • Is the value available at all when the record is created, or does it regularly arrive later?
  • Is there a fixed value list, or is free text unavoidable at this point?
  • Who decides when the value is disputed, and where does the binding version come from?
  • What happens concretely in the document if the field stays empty?
  • Does the value already exist elsewhere and can it be taken over instead of entered again?

Standardising spellings: rules instead of taste

Inconsistent spellings arise wherever free text is possible. The unit of measure appears written out, abbreviated and in capitals. The legal form sits before or after the company name. The district appears in brackets after the town or not at all. Phone numbers show up with a country code, with a zero in brackets, or with hyphens in varying places. A person can read all of that; no sorting, no search and no automatic matching can. The first step is therefore not clean-up but a short, binding rule for each affected field.

Spelling rules and checks at entry
Field           | Rule                                      | Check at entry
----------------|-------------------------------------------|--------------------------
Company name    | legal form written out and at the end     | search on the first five characters
Street          | house number in the same field, suffix    | cross-check with post code and town
                | directly attached                         |
Town            | without district in brackets              | value list per post code
Phone           | country code with a plus sign, rest       | check the format automatically
                | without separators                        |
Unit of measure | exclusively from the value list           | free text disabled technically
Article number  | fixed pattern, no special characters      | pattern stored in the system
Product group   | exactly one group per article             | mandatory field, no catch-all group

# rules that exist only on paper rarely survive a quarter
# every rule needs a place in the entry screen where it takes effect

Such rules only take effect once they act at the point of entry. A value list to choose from prevents deviations more reliably than any training session. A search that automatically looks for similar records before a new one is created prevents duplicates more effectively than a note during onboarding. Both can be configured in common systems and usually cost less effort than cleaning up the same errors later. Whatever cannot be enforced technically belongs in a short reference sheet that is available at the workplace and used when new staff are trained.

Numbers without meaning

Speaking numbers that encode a product group or a region become wrong as soon as the assignment changes. Sequential numbers without special characters are unremarkable and last longer. Deleted numbers are not reissued, otherwise old documents point at new objects.

Value lists instead of free text

Units of measure, countries, payment methods, tax codes and product groups come from a maintained list. Free text at these points is one of the most common reasons why a report falls apart into catch-all positions.

Search before creating

Before any new record, a search runs on number, start of the name and post code. The entry screen should offer that search itself rather than leaving it to the diligence of a person working under time pressure.

Make urgent cases visible

If a record may be created without mandatory fields under time pressure, that is exactly where tomorrow's legacy data grows. Better than an exception is a status on the record that demands completion and shows up in a working list.

Who maintains which field: responsibility, rights, control

Deciding on the leading system also means deciding on the role. A system is a place, not a person, and data quality comes from people who have been given a task. A useful split covers four rights: create, change, check, block. For most fields creating and changing coincide and sit with the department that needs the value. For a few fields a second person is appropriate, particularly for bank details, terms and discount levels, because an error there moves money immediately.

For access rights the principle applies of granting only as many rights as the respective task requires (German Federal Office for Information Security). In practice that means a technical account used for a transfer needs no right to change master data if it is only supposed to read. Equally, not everyone in sales needs the right to overwrite prices. The opposite mistake exists as well: if rights are set so tightly that every correction requires a request, side lists appear outside the system and the data ages again. The split should therefore be agreed with the departments and reviewed after a few weeks.

  • Creation: who may create a new customer, article or supplier at all, and in which system does that happen?
  • Change: who may change which field, and which fields are write-protected in the downstream systems?
  • Control: which changes need a second person, for example bank details, credit limit or discount level?
  • Blocking: who decides on blocking instead of deleting, so that older documents keep their reference?
  • Evidence: is it logged who changed which field and when, and can that log be evaluated?
  • Cover: who takes over maintenance during absences, so records do not stay incomplete for weeks?

Legally the topic touches several areas at once: the traceability of business transactions and the retention of documents, data protection for personal details such as named contacts and their contact data, and the question of how long records that are no longer needed may be kept. A written approach helps that describes which records are blocked rather than deleted and when a deletion takes place. This article does not replace legal advice; assessing your specific case belongs in professional review, particularly where retention obligations and deletion rights meet.

Cleaning up during operations: five steps

A clean-up that halts operations does not happen. The suggestion of working through the whole data set over a weekend founders on the volume and on the fact that the people with the necessary knowledge are needed in daily business at exactly that time. What works is an approach in waves that starts with the part of the data actually in use and lets the rest rest. In many companies a considerable share of records covers customers and articles without any movement in recent years; those do not have to be touched first, and some never do.

Field catalogue, mandatory fields, spelling rules and the leading system per field are written down and agreed with the departments. Without that basis, corrections follow personal taste, and six months later the same data set stands there with different deviations.

The most important point about this approach is the sequence. Rules first, clean-up second. Anyone tidying the existing data without securing new entries works against an inflow that is faster than their own progress. Conversely, a secured entry process takes effect before the legacy data has been touched at all, because the records created each day comply with the rules from then on. To estimate the effort it has proven useful to fully clean a sample of around one hundred records and extrapolate the time required (project experience).

A cleaned data set without rules for new entries is an interim state with an expiry date, not a result.

Basic rule from master data projects

Duplicates and legacy records: merge rather than delete

Duplicates are records created twice for the same customer, article or supplier. They appear when the search before creation is missing, or when a small difference in spelling means the existing record is not found. A simple comparison is rarely enough for detection: what works is a match on normalised values, where capitalisation, spaces, legal form suffixes and special characters are standardised before the comparison. The result is a list of candidates, not a list of facts.

Deciding whether two records mean the same partner belongs in the department and not in an automatic procedure. Two sites of one corporate group can be held under similar names and still belong apart, because they order separately and pay separately. If they are merged automatically, assignments are lost that daily business relies on. The following rules have proven useful for merging and apply equally to customers, suppliers and articles.

  1. Run the match on normalised values and treat the result as a candidate list, not as a decision.
  2. Have candidates confirmed by the responsible department, looking at ordering behaviour and payment route rather than at the name alone.
  3. Determine the surviving record: usually the one with the most documents and the most complete maintenance, not the one with the lowest number.
  4. Take missing values from the record being retired before it is blocked, so that no detail is lost.
  5. Block the retired record rather than deleting it, so older documents keep their reference and queries remain answerable.
  6. Store the old number as an additional search term and log every merge with date, person and reason.

The article master calls for extra caution, because behind an article number sit stock levels, prices, open orders and in some companies serial numbers or batches. Merging without agreement from the warehouse shifts stock and creates differences that surface during stocktaking and are then laborious to resolve. A workable route is to block the article being retired for new use, let the stock run out in an orderly way and retire the number only afterwards. The same principle applies to records from replaced legacy systems: what is needed is taken over, the rest stays in the archive rather than in the productive data.

Anchoring maintenance for the long run

Master data quality is not a state but an ongoing task with a small yet regular effort. After a clean-up, many companies get by with a short monthly review in which a few figures are looked at and open points are assigned. What matters is that this review has a fixed date and a named person. Without both it disappears into daily business, and after two or three years the same data set is back where the clean-up started.

A few figures, monthly

Share of empty mandatory fields, number of open duplicate candidates, number of records in completion status. Three figures are enough to see whether quality holds or slowly slips.

A named owner per data type

One person per data type decides queries and keeps the rules current. Without a name, maintenance stays a side matter and contested cases sit undecided in the data.

Rules where the work happens

The field catalogue belongs in the entry screen and in onboarding, not in a folder. A short, maintained version works more reliably than an extensive document with no link to daily practice.

Record the field catalogue, the responsibilities and the rules in one place that remains findable even when the person carrying the knowledge is no longer with the company. Such a description is at the same time a building block of your process documentation and pays off at every onboarding, every audit and every system change. It does not have to be extensive: a few pages of tables work better than a manual that no longer matches reality after a year.

A system change is the moment when the quality of the master data pays off or takes revenge. What exists is what gets migrated, and a migration is the worst time for a first clean-up, because deadline pressure and business clarification collide. Anyone planning a change should sort out the master data one or two quarters beforehand and start the migration with data that follows rules (project experience). The effort arises in either case; beforehand it can be planned, during the changeover it cannot.

This article is based on data from: the Federal Statistical Office of Germany, the German Federal Office for Information Security, the European Commission and our own project experience.

Related Articles

Data & documents

Finding and merging duplicates: customer and product records

Duplicate customer and product records: how they arise, how similarity matching and address normalisation find them, and how to merge without losing history.

14 min read
Process analysis & assessment

Duplicate data entry: three routes out of retyping

Typing the same data again and again: why duplicate entry arises, what it costs and the three routes out: a leading system, an interface, or file transfer.

13 min read
Practice & rollout

Replacing the grown spreadsheet: when and how it pays off

When a spreadsheet becomes a shadow core system: five weak points, three decision questions and a transition path with a parallel run instead of a standstill.

14 min read