Skip to content
Data & documents

Document management: from folders to a searchable archive

From a folder of PDF files to a searchable archive: indexing, linking to the case, versions, access rights, retention and how to handle the existing backlog.

14 min read DokumentenmanagementVerschlagwortungAufbewahrungVersionierungZugriffsrechte

In most mid-size companies a document can be found — just not necessarily quickly, and not necessarily in the version that currently applies. The route to it runs through a folder tree that grew over the years, through individual mailboxes and through the filing cabinet in the corridor. As long as the company is small and few people file anything, that works. With more cases, more people involved and more obligations it tips over: the folder path becomes longer than the file name, the same invoice sits in two versions in two places, and when old records are cleared out nobody knows what is actually there. This article describes the move from a folder full of files to a searchable document archive: which decision between filing structure and indexing is really on the table, why linking a document to its case matters more than any search function, how versions, access rights, retention and deletion hang together, what a system has to deliver and what folder structures still achieve. And it describes the point at which most projects run aground: the existing backlog. The wider context is set out on the page document digitisation.

Key takeaways

  • The real decision is not folder versus system but path versus attribute: as long as the order lives in the folder path alone, every document belongs to exactly one view, and multiple membership can only be represented through copies.
  • Five attributes carry the daily work: document type, the case it belongs to, the parties involved, document and receipt date, and version status. Companies that record these five reliably need far less search functionality than expected.
  • Linking a document to its case matters more than the full-text index: a document attached to an order, a customer or a machine is found through the case, even when nobody can recall the right search term.
  • Retention and deletion are two sides of one mechanism: the document type yields the period, the period yields the point at which someone checks and either extends or removes — by hand, in a folder tree, few organisations keep that up.
  • The existing backlog is rarely migrated in full: a cut-off date for new documents, a limited retrospective capture of open cases and loading older files on demand keep the effort contained while delivering the benefit early.

When a folder tree reaches its limit

A folder tree is not a poor solution, it is a very old one. It represents exactly one order, namely the one that seemed plausible when the first subfolder was created. As long as everyone involved carries the same order in their head, it works well. The break comes when a document belongs to two orders at once: a supplier invoice is also evidence for a customer order, a contract belongs to the customer and to the contract type, an inspection report to the machine and to the year of inspection. In a tree this can only be solved through copies, and from the first copy onwards two versions exist, one of which will eventually be out of date.

The warning signs look the same in almost every company. Folder paths become so deep that dialogues truncate them. File names carry suffixes such as final, final2 or copy of. Important attachments sit in mailboxes rather than on the shared drive because that route was quicker. After a colleague leaves, the team spends weeks looking for records that only she could reliably place. And when the tax adviser asks a question or an audit comes up, things get hectic, because everything is there but nobody can demonstrate that it is complete.

The effort rarely sits in the individual search but in the sum of small detours: asking which version applies, requesting a document a second time because the first cannot be found, or creating a filing location twice because the existing one was not located. Individually these detours are too small to be reported, and together they are large enough to justify a project (project experience). Anyone wanting to quantify them should measure them the way a process analysis measures other workflows: on real cases rather than on estimates from memory.

Three terms that are frequently mixed up

Filing means a document has a fixed place and can be found again. Document management additionally means every document carries attributes, a version status, access rights and an owner. Audit-proof archiving goes further still: a document is stored unalterably, access is logged, and it stays available over the statutory period. The three levels build on each other but cost very different amounts. Many companies need the third level only for part of their documents rather than for the whole archive.

Filing structure or indexing: the decision that really matters

Behind the question of the right system sits a simpler question: what does the order hang on? With a filing structure it hangs on the path. The place in the tree states what the document is about, and the document sits in exactly one place. With indexing, the order hangs on attributes stored on the document itself. The storage location loses its meaning, and the same document can be reached through several routes: through the customer, the order, the document type or the period.

The two approaches are not mutually exclusive. In practice a flat base structure with few clearly named areas plus attributes per document works well. What rarely works is trying to rescue a finely branched folder structure after the fact with search functions. A full-text index over an unsorted archive finds a great deal, but it does not answer whether the version found is the version that applies. That question is what costs the most time in daily work.

CriterionFolder structureIndexingBoth combined
Effort to startLow, already in placeModerate, define attributesModerate to higher
Finding things againThrough the known pathThrough several attributesThrough path and attributes
Multiple membershipOnly through copiesOne document, several viewsOne document, several views
Maintenance effortGrows with the depthRequires disciplined captureRules per document type
Retention and deletionBy hand, easily forgottenDerivable from attributesDerivable from attributes
Best suited toFew document typesMany similar recordsA mixed archive

The choice of attributes decides how useful the archive later becomes. Too few attributes produce a search that returns too many hits. Too many attributes lead to fields being skipped during capture, leaving the archive incomplete. A workable start is a small set of mandatory attributes per document type, supplemented by optional entries maintained only where they genuinely help. Whatever is mandatory should, where possible, be filled automatically, for instance from the case in which the document is created.

Documents in a company are not created for their own sake but for a case: an order, a customer, a machine, an employment relationship or a delivery. That very membership is lost the moment the document is filed as a file. It then survives in the file name, in the folder path or in the head of the person who filed it. An archive in which every document knows its case reverses the search: instead of looking for the document, you open the case and see everything belonging to it — quotation, order confirmation, delivery note, inspection report, invoice, correspondence.

This reversal is the practical core of document management. It also relieves colleagues who rarely work with the archive, because they no longer have to guess search terms. And it makes gaps visible: if a completed order is missing the signed delivery note, that stands out when looking at the case, whereas it goes unnoticed in a folder tree. The prerequisite is that the case has a stable identifier used across the whole company — the same number in the merchandise management system, in the archive and in correspondence.

  • Document type: quotation, order confirmation, delivery note, incoming invoice, contract, report, certificate
  • Case: the identifier the document belongs to in business terms, such as an order or asset number
  • Parties: customer, supplier or the person concerned, each via the number from the leading system
  • Document date and receipt date: recorded separately, because periods hang on one and workflows on the other
  • Version status and approval: which version applies and who approved it
  • Retention class: the rule from which the period follows, not the date itself
  • Confidentiality: open within the company, limited to one department, or specially protected

These attributes do not all have to be typed in. Where documents are produced out of a system, such as quotations and invoices from merchandise management, case, parties and date can be carried across directly. For incoming post, a pre-fill based on sender and document type that the capturing person merely confirms pays off quickly. The fewer keystrokes capture requires, the more complete the archive becomes — that is the most reliable rule in this field (project experience).

Versions: which one applies

Versions are where a folder tree hits the wall fastest. As soon as a document is revised repeatedly, file names sprout suffixes, parallel versions appear in mailboxes, and the recurring question arises of which version was actually sent. An archive solves this by attaching the versions to one object: there is one document with a history, not several files with similar names. Whoever opens the history sees the time, the person and ideally a short note on what changed.

The separation between working drafts and approved versions matters. There may be many working drafts; they concern only the person editing. Exactly one version is approved, and that is the one that goes out. Where this separation is missing, the familiar uncertainty before sending arises, which usually ends in a query to a manager. For documents that leave the company it should therefore be settled who approves and how approval is recognisable.

From file name to attribute
# So far: the order lives in the name
2026-06-24_quotation_A-2026-0412_v3_final.pdf

# From now on: the name is just decoration
Document type:   Quotation
Case:            A-2026-0412
Parties:         customer number 10457
Document date:   2026-06-24
Receipt date:    not applicable, produced in house
Version:         3, approved by sales management
Retention:       commercial and business letter
Confidentiality: open within the company

# Rule: attributes that come from the case are carried across
# and never typed in again

A naming convention still makes sense as long as documents leave the archive and travel as attachments. The difference is that the file name is then a courtesy towards the recipient and no longer a place where data is held. Companies that make this shift cleanly stop arguing about folder names and start discussing attributes instead — the more productive discussion, because it follows tasks rather than habits.

Access rights: reading, changing and deleting are three separate questions

On a shared drive, rights usually apply per folder. That is coarse and leads to two familiar patterns: either almost the entire team can see almost everything, or special folders appear whose permissions nobody can survey any more. In a document archive, rights hang on attributes instead: on the document type, on the confidentiality level, on the department it belongs to. That allows rules to be written which also apply to future documents without anyone having to remember them.

The separation of read, change and delete permissions is particularly relevant. Many companies effectively grant all three together because the shared drive barely allows anything else. For personnel records, contracts and certificates that is inappropriate. It also helps to treat deletion not as an individual permission but as the result of a period: what may be deleted follows from the retention rule, not from the judgement of whoever happens to be tidying up.

Roles instead of individual grants

Rights are granted to roles such as accounting, order processing, personnel or management. People receive roles. When a responsibility changes, the role changes rather than a list of individual grants that nobody knows in full.

Confidentiality as an attribute

The protection level hangs on the document, not on the storage location. A personnel record therefore stays protected even when it belongs to a case other roles may view, and moving it does not change its visibility by accident.

Logging with a purpose

Logging access is worthwhile only when it is clear who evaluates the log, when and on what occasion. Without that purpose, a body of data about your own staff accumulates that itself needs justifying.

Settle codetermination early

Systems that record access and edits in a personally identifiable way touch codetermination. Where a works council exists, it belongs in the conversation early. Assessing the specific case remains a matter for professional and legal review.

A simple check before rollout: take five typical document types and answer three questions for each — who may read it, who may change it, who decides on deletion. If the answers differ around the table, that is where the real need for clarification lies, not in the software selection.

Retention and deletion: two obligations, one mechanism

Two requirements pull in opposite directions. Tax and commercial law require certain records to remain available for years. Data protection requires that personal data is not stored longer than the purpose demands. Satisfying both at once is hard in a folder tree, because neither the period nor the purpose is recorded on the document. In an archive with attributes it becomes a derivation: document type and document date yield the period, and the period yields the point at which the case is reviewed.

Orders of magnitude for retention

Commercial books, inventories, annual accounts and management reports are subject to a retention period of ten years, accounting vouchers to eight years, and commercial and business letters received and sent to six years (German Fiscal Code). The period starts at the end of the calendar year in which the record was created. For the traceability of digital workflows, process documentation is additionally expected (German Federal Ministry of Finance). Which period applies in an individual case and which special rules bite belongs in a conversation with your tax adviser; this article does not replace legal advice.

In practice this means every document type receives a retention class. The rule hangs on the class, not on the individual document. If the legal position changes, the rule is adjusted and takes effect across the whole archive. When a period expires, no automatic deletion follows but a list for review, because there are regularly reasons to keep a document longer: ongoing proceedings, an open warranty, a case that has not been closed. Skipping this intermediate step means losing records that were still needed sooner or later.

Deletion is a function, not a clear-out day

In many companies deletion happens as an event: someone tidies up a drive because storage is running short. That is the worst possible moment and the worst possible trigger. The sensible route is the reverse: the period sits on the document, the system presents a review list when it expires, a named person decides, and the decision is logged. An irregular event becomes a recurring workflow with evidence — and the evidence is what counts when questions arise.

What a system has to deliver and what folders still achieve

Not every company needs a document management system. Where few document types occur, allocation is unambiguous and the archive stays manageable, an orderly filing structure with clear rules carries further than the vendor landscape suggests. Folder structures still achieve plenty: a coarse split by area, a place for completed years, a defined location for templates and a structure readable by people without any additional software. The switch pays off where multiple membership, retention periods, access rights or evidence obligations come into play.

If a system is on the table, only a handful of properties decide the outcome in daily work. The list below describes requirements, not products. It works as a review grid for quotations and protects against demonstrations in which a lot is shown and little is answered. The last point matters most: an archive from which documents can be extracted together with their attributes stays available even if the system is replaced later — the same consideration that applies when replacing legacy systems.

  • Capture without detours: scanning, dragging in from the mailbox, taking documents over from the producing system
  • Text recognition across image files, so that scanned records can be found in full text as well
  • Attributes per document type with mandatory fields and automatic pre-fill from the case
  • Links to cases and parties, with a jump from the case to every document belonging to it
  • Version history with approval, visible to everyone entitled to use the document
  • Rights per role and confidentiality level, separated into reading, changing and deleting
  • Retention classes with a review list when a period expires rather than silent deletion
  • Export of the archive with all attributes in an open format, at any time and without extra charges

The decisive question during selection is not what the system can do in total, but how many keystrokes it takes to capture an ordinary incoming record. Everything else is decided later; this point is decided on day one.

Review rule from rollout projects

One more point: a system that sits beside the existing workflows will be bypassed. It should appear where the work happens — in the case, in the order, in the inbox. If filing is an extra step at the end of the day, it is left undone. If it is part of the case, it happens along the way. This placement belongs to process documentation and should be settled before the selection, not after it.

The existing backlog: why migration should stay limited

The most common reason projects fail is not the technology but the wish for completeness. Deciding to capture and index the entire existing archive ties staff up for months in work with no visible progress. The arithmetic turns unfavourable quickly: an archive of several tens of thousands of files needs only a few minutes of review and allocation per document to add up to person-months. At the same time a large share of those documents will never be opened again after the move.

The workable answer is a cut-off date. From an agreed date, new documents are created exclusively in the new archive. The old folder tree stays available for reading and is frozen: nothing is added and nothing is changed. From day one the order is unambiguous, because the question is no longer where does it sit but what date is it from. For the transition a short rule that everyone in the company can remember is enough.

Retrospective capture is then deliberately limited. What gets captured is whatever belongs to open cases, whatever is needed regularly and whatever has to remain legally available. Everything else follows on demand: when an old document is needed, it is captured once while being retrieved and from then on forms part of the archive. Loading on demand spreads the effort across many shoulders and many months — and it captures precisely those documents that actually matter (project experience).

A date from which new documents are created only in the new archive. The cut-off should fall on a quiet point in the year rather than on the annual accounts or the peak season, and it should be announced at least two weeks in advance.

There is one exception to this restraint. If the old archive is technically at risk, for instance because it sits on ageing equipment or in a format that will no longer be readable, then migrating it is a question of availability rather than convenience. In that case it pays to separate migration from indexing: secure the material first, then add attributes step by step.

Rolling out in steps rather than in one leap

A start works better when it begins with one document type rather than with the whole company. Incoming invoices are a good candidate: they arrive regularly, look alike, clearly belong to a case, and their retention is settled. Once that workflow stands and the people involved have mastered it, contracts, certificates and correspondence follow. Every further document type is then an extension of familiar rules rather than a new project.

A short check on the first weeks belongs to the rollout. Not as reporting but as an early warning: how many documents arrive, how many are allocated to a case, how many stay unresolved, and for what reason. Unresolved cases are the most interesting figure, because they show where attributes are unclear or pre-fills are missing. After a few weeks that number usually falls noticeably, provided the feedback is acted upon.

Terminal
$ # Check the daily overview of incoming records
$ docimport --run 2026-06-24 --status
$ 128 files captured, 121 allocated to a case, 7 unresolved
$ # List unresolved items with the reason
$ docimport --unresolved --details
$ 4x case not recognised, 2x document type unclear, 1x unreadable scan

Just as important as the technology is training on real cases. A session in which staff capture their own documents from the current week has a more lasting effect than a demonstration on sample data. Two short sessions two weeks apart achieve more than one long session at the start, because the second session is where the questions that arose in daily work come up. How such a transition is structured is described on the page training and rollout.

Finally, the new archive needs an owner. Attributes go stale, document types are added, periods change, new staff need an introduction. A named person with a fixed slot each month is usually enough in companies of this size. Without that ownership any order decays, and in a few years the company stands where it began — only with a different tool.

This article is based on data from: the German Fiscal Code, the German Federal Ministry of Finance, and project experience from rollout projects in mid-size companies.

Related Articles

Data & documents

Digital personnel files: access, retention, evidence

Which section of a personnel file carries which retention period, who may access it, what gets logged, and how inspection and access become routine cases.

18 min read
Law, security & funding

Writing process documentation for tax audits

Process documentation under the German GoBD rules: the four required parts, how detailed it must be, how to keep it current and what its absence can mean.

14 min read
Data & documents

Text recognition in practice: what it reads, what it guesses

Text recognition realistically assessed: clean sources versus carbon copies, stamps and handwriting, measuring quality, fields to extract, effort per type.

14 min read