What Your 10–20% Duplicate Rate Is Really Telling You

A 10–20% duplicate rate in your material master is an organisational symptom, not a data error. Learn to read it — and fix the root cause.

Data Duplication as an Organisational Health Indicator: What the 10–20% Figure Really Tells You

If somewhere between 10 and 20 per cent of the records in your material master are duplicates — and in most asset-intensive organisations, they are — why do you keep calling it a data problem? That question sounds pedantic until you sit with it for a moment. Data does not duplicate itself. Somewhere in your organisation, a person created a record for an item that already existed, and the system let them. Then it happened again. And again. Thousands of times, across years and sites, until duplication became a measurable percentage of your catalogue.

Read that way, a duplication rate is not a technical metric at all. It is a diagnostic reading — the organisational equivalent of a blood test result. And like any diagnostic reading, the number itself matters far less than what it reveals about the system that produced it.

Duplicate Records Are a Symptom, Not the Disease

Most material master data cleansing projects begin with the same framing: the data is "dirty", so we will clean it. Duplicates are identified, merged and retired; descriptions are standardised; classifications are corrected. Twelve to eighteen months later, the duplication rate is climbing again, and everyone quietly wonders why the investment did not stick.

The reason is straightforward. Duplication is a symptom, and the project treated the symptom. The disease — the set of organisational conditions that generated the duplicates in the first place — was never diagnosed, let alone addressed. No physician would treat a recurring fever by simply lowering the patient's temperature and sending them home. Yet that is precisely how many organisations approach their material master: reduce the number, declare success, and wait for the relapse.

The more useful move is to treat your duplication rate the way a doctor treats a symptom: as evidence. Every duplicate record is a small fossil of an organisational failure. It tells you exactly where, when and why your processes broke down — if you are willing to read it.

Contact Panemu

Three Organisational Failures Hiding Inside Your Duplicate Rate

When our cataloguers work through a client's material master data, the duplicates almost always trace back to three underlying conditions. They are worth naming precisely, because each demands a different remedy.

No quality gate at item creation. In many ERP, EAM and CMMS environments, creating a new material record is easier than finding an existing one. There is no mandated search step, no naming convention enforced at entry, no review before a record goes live. A storeperson under pressure to receive parts, or a maintenance planner racing to close a work order, will take the path of least resistance — and the path of least resistance is a new record. A duplication rate of 10–20 per cent is the physical trace of thousands of these moments where speed beat control, because control was never designed into the process.

Sites that do not talk to each other. In multi-site operations — mining, oil and gas, power generation, manufacturing — duplication clusters along organisational boundaries. Site A describes a bearing one way; Site B describes the same bearing another way; head office cannot tell they are the same item. This is not carelessness. It is the predictable result of sites operating as data silos, each with its own conventions, abbreviations and local knowledge. The duplicates are simply where the silo walls become visible in the data. If your duplication analysis shows the same manufacturer part number appearing under three different descriptions at three different sites, you have not found a data error. You have found an organisational structure problem, written in your inventory records.

A workforce that has given up on the system. This is the most telling failure, and the hardest to see from a dashboard. When users have searched for an item repeatedly and failed to find it — because descriptions are inconsistent, classifications are wrong, or the search simply does not work — they stop trusting the catalogue. From that point on, creating a new record is not laziness; it is a rational adaptation to a system that has failed them. Duplication becomes a workaround culture made permanent. And workaround cultures rarely confine themselves to the material master. Where you find one, you usually find shadow spreadsheets, informal procurement channels and tribal knowledge doing the work your systems were supposed to do.

What the Number Costs You While You Ignore It

If duplication were merely untidy, it could wait. It is not. Each of the three failures above compounds into hard commercial consequences, and the duplicates are the mechanism.

Consider working capital first. Every duplicate record is an opportunity to hold the same spare part twice — once under each identity. Inventory optimisation tools cannot consolidate stock they cannot recognise as identical, so safety stock is calculated separately for each record, and capital quietly accumulates on warehouse shelves. In organisations with tens of thousands of stock items, the value tied up in unrecognised duplicate holdings routinely runs into the millions.

Procurement suffers next. Duplicate records fragment purchase history, which weakens demand visibility and undermines supplier negotiations. Buyers order items that already sit in another warehouse under a different number. Expediting fees are paid for parts the organisation already owns. And because spend is scattered across duplicate identities, category managers negotiate from a picture of demand that is systematically understated.

Maintenance carries the operational risk. When a planner cannot confidently locate the correct part, work orders are delayed, kitting errors increase, and — in the worst cases — equipment sits idle while a part that exists on site is reordered from a supplier. In capital-intensive industries, where an hour of unplanned downtime can cost more than an entire cataloguing programme, this is not a data-quality footnote. It is a reliability problem.

None of these costs appears on a report labelled "duplication". They surface as inventory variance, procurement inefficiency and maintenance delay — which is exactly why the root cause so often escapes attention.

Contact Panemu todayy

Why Cleansing Alone Guarantees a Relapse

Here is the uncomfortable position this article is prepared to take: a data cleansing project that only cleanses data is a waste of money. Not because the cleansing is done badly, but because it changes nothing about the conditions that produced the duplicates. The quality gate is still missing. The sites still do not share conventions. The users still do not trust the search.

A credible remediation programme therefore has to work at two levels simultaneously. At the data level, it must resolve the existing backlog: identifying true duplicates (which requires genuine cataloguing expertise, because two records with different descriptions may be the same item, and two near-identical descriptions may be different items), standardising names and descriptions against a consistent convention, and classifying items to recognised standards such as UNSPSC or NATO codification so the catalogue becomes searchable in the first place.

At the organisational level, it must change how records come into existence. That means defined naming and description standards that every site follows; a governed item-creation workflow with a mandatory duplicate check before any new record is approved; clear ownership of the material master as a managed asset rather than an unattended by-product of ERP transactions; and enough attention to search usability that users can actually find what exists — because the fastest way to stop duplicate creation is to make finding easier than creating.

This is where professional cataloguing services earn their keep. At Panemu, our cataloguing team combines both levels in one engagement: experienced cataloguers perform the cleansing, naming, describing and classification work item by item, supported by purpose-built tooling — our SCS platform assists with data cleansing, naming, describing and classification, always under the guardianship of the cataloguing team rather than replacing their judgement. The deliverable is not just a cleaner dataset. It is a standardised, classified, governable material master, along with the conventions and governance recommendations your organisation needs to keep it that way. Clients in mining, oil and gas, power generation and manufacturing engage us precisely because the second part — preventing the relapse — is what protects the investment in the first part.

Reading Your Own Diagnostic

If you want to apply this thinking to your own organisation this week, you do not need a project. You need an honest look at three questions.

First, trace ten recent duplicates back to their creation. Who created them, under what pressure, and what would they have needed to find the existing record instead? The answers will show you where your quality gate should sit.

Second, map your duplicates against your organisational structure. If they cluster along site or business-unit boundaries, your problem is coordination, not competence — and no amount of cleansing will fix a silo.

Third, ask your storepeople and planners a blunt question: when you cannot find an item, do you assume it does not exist, or do you assume the search has failed you? If the answer is the latter, your duplication rate will keep climbing no matter what you do to the data, because your people have already stopped believing in the catalogue.

Contact Panemu now

Conclusion

A duplication rate of 10–20 per cent is not an embarrassment to be quietly cleaned up. It is one of the most honest pieces of management information your organisation produces — a direct measurement of missing controls, disconnected sites and eroded user trust, recorded automatically and without spin. Organisations that treat the number as the problem will cleanse, relapse and cleanse again. Organisations that treat the number as a symptom will fix the process, the governance and the searchability behind it — and find that the number takes care of itself.

The difference between the two is not budget or technology. It is diagnosis.

Ready to Read What Your Own Data Is Telling You?

Before your next data cleansing initiative, there is a simpler question worth asking: do you actually know what your duplication rate is — and what it is trying to tell you?

Because while the number sits unexamined, the consequences do not wait. Procurement decisions slow down. Inventory visibility narrows. Parts that already sit in a warehouse get bought again. Working capital stays locked on shelves, and skilled people spend their days hunting for information instead of acting on it.

And in most cases, the fault is not your ERP, and it is certainly not your team. It is the quality, governance and searchability of the material master data underneath them — inconsistent descriptions, unreliable classifications, and a catalogue nobody can confidently search.

That is why at Panemu, we help organisations understand the real condition of their material master data through a free consultation and data assessment — identifying hidden duplicates and quality issues, evaluating how your data got to where it is, and providing practical recommendations for a stronger procurement, maintenance and supply chain foundation.

Because better decisions start with better data.

Curious whether your material master is supporting your business — or quietly working against it?

Send us a sample of your material master data for a free assessment, or book a consultation with our cataloguing team at https://panemu.com/cataloguing-service.