Predictive Maintenance Field Frameworks: The Evidence Chain
Predictive Maintenance Part 4 of 12

Predictive Maintenance Is a Data-Quality Problem First

The programme rarely fails at the model. It fails at the join, at the failure code nobody filled in, and at the reading stored without what the machine was doing at the time — none of which raises an error, because a broken data chain still produces output.

Article cover: predictive maintenance is a data-quality problem first.

A predictive-maintenance programme almost never fails at the model.

It fails earlier and much more quietly. A vibration point that cannot be joined to a work order because the historian calls the machine P-101 and the CMMS calls it PUMP-101. Three years of corrective history whose most frequently recorded failure mode is “general repair”. A temperature trend with no record of what the machine was doing at the time, so nobody can say afterwards whether the rise was a developing fault or a heavier grade on a warm afternoon.

None of that raises an error. Dashboards populate, models train, output arrives the whole way down. A broken data chain is not silent — it is fluent.

So here is a position, deliberately strong. In predictive maintenance, data quality is a primary technical risk rather than a hygiene factor — in most programmes ahead of sensor selection, and ahead of model choice. If your own setbacks trace somewhere else, that is worth knowing too, and the last section is the instrument for finding out.

Four forces, and not one of them is carelessness

Nobody made the bad data. It accumulates from four structural forces that recur wherever records are left unmanaged.

ForceWhat it doesWhere the cure lives
Many authors, no dictionaryone machine described five ways by five systems and five shiftsa naming standard; a failure-code list
Purpose mismatchdata recorded to close a work order, not describe healthdefined event and reading models
Missing contextreadings stored without the load, speed and state that give them meaningan operating-state model
Silent decayassets renamed, sensors moved, standards drifting, no record of whengovernance, and a named owner

That table explains why the narrow pilot so often outperforms the broad rollout. A pilot’s data can be curated by hand; a programme’s can only be governed by structure, built in dependency order, because each later cure assumes the earlier ones.

What three years of work orders is actually worth

The CMMS history is the most valuable dataset a programme has and its dirtiest — the plant’s only first-hand account of how its own machines die, written in shorthand, under time pressure, by people whose job was fixing the machine rather than describing it. The U.S. National Institute of Standards and Technology took that seriously enough to propose a discipline for it, technical language processing, arguing that maintenance text (“chkd brg, snd OK, adj”) defeats tools trained on prose and needs methods, dictionaries and human-in-the-loop workflows of its own (Brundage, Sexton et al., 2021).

Excavating it is a bounded, one-time job in five passes. Join every work order to a register asset. Classify corrective work apart from PMs, projects and admin. Code the failure modes, keyword rules proposing and a human confirming samples until the rules earn trust. Date the events, because failure, report and completion dates diverge. Cost them, because a baseline needs money attached.

Two disciplines govern the coding pass. Rules assist, humans decide. And be honest about residue: forcing a code onto a record that will not yield one corrupts the statistics. A record coded “unknown” is data. A record guessed is contamination.

At the Meridian plant — an illustrative composite, not a client — the excavation ran three weeks part-time across 1,940 work orders from three years. 88% joined cleanly to register assets. 71% of those were corrective rather than PM, project or admin. Rules-plus-review coding left 12% of that corrective subset marked “unknown”, logged without shame. The thousand-odd coded records remaining became 214 dated, priced events once narrowed to nine asset classes and de-duplicated, a failure with three follow-up orders counting once.

A five-stage funnel labelled illustrative composite: 1,940 work orders over three years, 88 percent joining cleanly to register assets, 71 percent of those corrective, 12 percent of that corrective subset left marked unknown, and 214 dated and priced events remaining after narrowing to nine asset classes and de-duplicating.
The yield is not the point — the arithmetic is. Every stage is a number a plant can produce about its own archive, and the unknowns are kept rather than filled in.

That 214 is not a benchmark and will not transfer. What transfers is the shape: an archive, a join rate, a corrective fraction, an honest residue, and a much smaller set of events that can carry a baseline.

One plant, one language

In the early 1970s the German power industry faced a naming crisis of its own making: every utility, often every station, labelled identical equipment differently. Its answer, the Kraftwerk-Kennzeichensystem (KKS), was a hierarchical designation system in which one code names plant, system, equipment unit and component. Its modern descendant, RDS-PP, follows the reference-designation structures of ISO/IEC 81346.

Schemes of that kind exist, are long-established and are documented. That is all the history supports — it does not make any scheme you write one of them, and adopting the pattern is not conformity with anything. What is borrowable is the discipline underneath, which fits on one page.

  • Address from the hierarchy: every tag opens with its functional-location path. A tag whose machine cannot be found from its name is an orphan at birth.
  • Semantics after the address: quantity, technique and point, from a controlled list — not from whoever configured the gateway.
  • Slot and equipment kept apart: condition history travels with the physical machine, where the physics lives; limits belong to the position.
  • Aliases mapped, never multiplied: historian, CMMS and vendor names grandfathered into a table with one master name each.

A standard that lives in a document dies quickly; a standard that lives in the systems endures. In order of leverage: the register and CMMS as master source, a gateway configuration that refuses to publish an unmapped tag, the protocol layer carrying semantics in transit, and a procurement data-delivery clause obliging new vendors to speak the site’s names — contract language, so drafted with procurement and legal, not by the reliability team alone.

Six ways to be wrong about your own data

“Good data” is an adjective until it is measured. ISO/IEC 25012 is one published attempt to name the dimensions; it supplies the left column, and the plant questions beside it are ordinary reliability questions, not a claim of conformity.

DimensionThe plant question it turns into
Completenessare the readings and the codes actually there?
Accuracydo the values reflect reality — calibration, range, physics checks?
Consistencydo the register, historian and CMMS agree what a machine is called?
Timelinessdid it arrive in time to matter, and how much was backfilled late?
Traceabilitycan each value name its origin — point, instrument, state?
Validitydoes each record conform to its own schema and units?

Three rules keep the scorecard useful rather than decorative. Score per consumer, not per warehouse: a threshold running on eleven tags needs completeness on those eleven, not a plant-wide average that hides them. Alarm on trend, not level: route completeness sliding 98 → 91 → 84% is a programme dying, visible well before anyone names it. And every red cell gets an owner, a date and the action that clears it, or the scorecard becomes wallpaper with percentages on it.

The failures worth fearing are the silent ones: not a wrong number but a missing one, invisible because every number that survived looked fine.

The exercise: audit readiness before ambition

The plant-wide audit is five questions on one page. Here is the one-asset version — an afternoon, no capital, four questions, each owed evidence rather than recollection.

  1. 01
    Join

    Take one critical asset. Can every condition source on it be joined to its register entry, by ID, today? Show the join, do not assert it.

  2. 02
    Code

    Pull twelve months of its work orders. What fraction carries a failure code more useful than 'general repair'? Read ten in full.

  3. 03
    Context

    For one condition tag, can you recover what the machine was doing — load, speed, state — at the time of any reading you pick?

  4. 04
    Own

    Write a name against each of the three answers. If it is the same name three times, you have found the finding.

At Meridian (illustrative), the plant-wide version took a planner and an engineer four days: register joins sound, 61% of three years’ work orders carrying no failure code beyond “general repair”, four of eleven condition-relevant historian tags failing the compression audit, load state recoverable for compressors and pumps but not conveyors — and every answer with the same owner, which was itself the problem.

What will probably happen is that questions one and four answer in minutes and questions two and three do not. That asymmetry is the result: identity is usually in better shape than description, so the cheapest next move is coding discipline rather than a sensor. A pilot that books two weeks for modelling and two days for data has almost certainly inverted the real split.

Which of those four questions can you answer with evidence rather than recollection — and who owns the one you cannot?

The four structural forces, the five-pass excavation method, the naming-standard rules and the six-dimension scorecard are from Predictive Maintenance: Practitioner Reference Frameworks and Planning Guide (Part 7: Data Foundations for Predictive Maintenance).

Lokesh Chennuru
Lokesh Chennuru
Industry Digits Author

Lokesh Chennuru writes Industry Digits field notes for industrial decision makers, focused on automation, IIoT, condition monitoring, predictive maintenance, and industrial AI.

Connect on LinkedIn
Frequently asked

Questions industrial leaders ask about this

Why does predictive maintenance fail on data quality rather than on the model?

Because a broken data chain still produces output. Dashboards populate and models train whether or not a condition reading can be joined to the machine it came from, whether or not the work order describing the last failure carries a usable code, and whether or not anyone recorded what the machine was doing at the time. Nothing raises an error, so the problem survives until somebody counts.

What is a data-readiness audit?

Five questions, answered with evidence rather than recollection, before any analytics is funded. Can every condition source be joined to the asset register? What fraction of work orders carries a usable failure code? Do the historian tags that matter survive a compression audit? Is operating state recoverable for the assets that matter? And who, by name, owns each of those answers?

Is the $3.1 trillion cost-of-bad-data figure usable in a business case?

No. Thomas Redman published it in a Harvard Business Review column in 2016, attributing it to IBM. It is a decade-old modelled estimate for the U.S. economy as a whole, produced by a method that was never published, and it is not divisible into a figure for any plant, sector or maintenance function. It illustrates the shape of a problem and supports no local number.

Should uncoded maintenance records be forced into a failure code?

No. Rules propose codes and a human confirms them against samples of the actual work-order text; whatever will not yield a code is marked unknown and logged. A record coded unknown is data. A record guessed is contamination, and it contaminates every statistic computed downstream of it.

What can a data-quality scorecard not tell you?

It cannot tell you a programme will work, and it does not authorise any maintenance decision. It reports six measured dimensions — completeness, accuracy, consistency, timeliness, traceability, validity — per consumer, so a red cell names a specific gap with an owner and a next action. A green card means the data is describable, not that the evidence behind any alert is sound.

Go deeper

Predictive Maintenance — Practitioner Reference Frameworks and Planning Guide

The twelve-part reference this series draws on: foundations and the value case, asset criticality and strategy, failure modes and degradation, the monitoring technologies, asset-class playbooks, sensors and IIoT architecture, data foundations, signal processing, analytics and prediction models, alerts and diagnosis, work management and CMMS integration, and pilot execution through rollout and governance — 126 sections with 46 technical figures.