Build the Data Chain Backwards From the Decision
Eight links carry a plant reading from a physical condition to somebody's judgement, and value is created at the last one only. Built forwards from the sensors, the chain produces a very complete record of nobody looking at anything.
The uncomfortable arithmetic of plant data is that collecting it has become cheap, and that is precisely the problem.
A sensor that once justified a week of engineering can be clamped on in an hour. A historian that once meant a server room now means a subscription. What has not become cheaper is attention — the engineering hours to give a signal context, the operations hours to look at it, and the management will to act on what it shows.
So the constraint moved. It is no longer collection. It has been the decision at the end of the chain for years, and most architecture reviews still start at the historian, which is to say they start after every decision that mattered has already been made.
Eight links, one of which pays
Every industrial data project moves along the same chain whether anyone draws it or not.
Process → sensor → signal → acquisition → transport → storage → context → decision.
A physical process produces a condition. A sensor turns it into a signal. An acquisition layer — a controller, an RTU, a meter, a gateway — turns the signal into a named, timestamped value. A network moves it. A store keeps it. A context layer makes it joinable and meaningful. A consumption layer puts it in front of somebody.
Value appears at the last link only. Everything before it is cost, and each link has its own currency: engineering at the sensor, hardware at the signal, configuration at acquisition, bandwidth in transport, licence at storage, modelling in context. Attention is spent at the end, and it is the one nobody budgets.
This is an editorial organising model, not a standard and not a map to one. Its whole job is to make the hand-offs visible, so that a disciplined question can be asked at each of them.
| Stage | The review question | What a missing answer costs later |
|---|---|---|
| Process | Which physical condition matters, and to which decision? | a measurement that never had a job |
| Sensor | What can be measured, at what accuracy, cost and installation burden? | a datasheet bought instead of a duty |
| Signal | How does the measurement become an electrical or digital value, and what corrupts it on the way? | noise mistaken for process behaviour |
| Acquisition | Which device turns signals into data, with what timestamps, rates and buffering? | properties the chain can never add back |
| Transport | Which networks and protocols move it, with what capacity, latency and failure behaviour? | the worst minute discovered in production |
| Storage | Where does it rest, at what resolution, for how long — and what did compression do to it? | an event erased by a setting made years ago |
| Context | What names, models, units and quality codes make it joinable? | two systems, one pump, no join |
| Decision | Who consumes it, in what display, feeding which decision — and did behaviour change? | a programme with output and no evidence |
Two concerns cut across every stage rather than living at one of them. Security, because every link is also an attack surface. Ownership, because every link needs a named owner for its configuration, its failure alarms and its changes — and “the data project” is not a phone number at two in the morning.
The direction of travel is the whole method
Walked forwards, the chain is a shopping list. Machine first, catalogue open, and every question defaults to more, better, everywhere. The economics drown quietly, because each individual purchase is defensible and the total is not.
Walked backwards, the same chain becomes tractable. Start at the decision. Name it — a recurring decision with a role attached, not “visibility”. Name what value or trend would change it. Then walk to the physics: what effect carries the earliest useful evidence for this decision, what accuracy this decision actually needs, and where a sensor would have to live to see the effect at all.
The accuracy question is where backwards-thinking pays fastest. A trend feeding a scheduling decision tolerates an offset error that would ruin a custody-transfer measurement. Walked forwards, both get specified to the tighter number, and the difference is real money spent on precision no decision will ever consume.
Coverage is a portfolio, not a percentage
Once individual measurements can be justified, the plant-level question arrives: of everything that could be measured, what should be?
The useful frame is a portfolio sorted by decision value, not a coverage percentage. A reliability review leaves a ranked list of failure modes with consequences attached. An energy review leaves a ranked list of loads. A quality review leaves process parameters with scrap cost against them. Each list is a queue of named decisions. The sensing programme works down the queues in value order and stops where the chain cost exceeds the decision value.
At the Meridian plant — an invented composite used throughout the underlying reference, not a client — that queue has four funded rows and a fifth that reads everything else on the wish list: queued until a decision and owner exist.
The last row is the discipline, and it works only if the waiting list is public. Anyone can promote a candidate by naming the decision, the owner and the value. What nobody gets to do is promote a candidate because the sensor is cheap. Cheap sensors still buy engineering time, network capacity, storage and attention, and the expensive links are all downstream of the cheap one.
An exercise that costs an afternoon
Take a data project already funded — one where hardware has been chosen and something is being built.
Write the eight links down the left of a page. Against each, fill three columns from what is genuinely written down today, not from what people believe: who owns this link, what it cost, and what it would look like if it failed. A stranger should be able to read the answers.
Then do the one that hurts. In the Decision row, write the sentence: “When this value crosses ___, [named role] will ___ instead of ___.”
Three findings recur, and none of them costs capital to discover.
- 01 The decision sentence will not finish
The value is named, the role is vague, and the alternative action is missing. That is not a sensor problem and no amount of storage will fix it.
- 02 Two or three links have the same owner
Usually the person who built the thing. A chain owned end-to-end by one enthusiast is a chain with one failure mode, and it is not technical.
- 03 Nobody can describe silent failure
Ask what a stuck collector, a full buffer or a drifted clock would look like on the dashboard. If the answer is that it would look normal, that is the finding.
The point of writing the chain out is not documentation. It is that cost, quality and ownership become visible in the same picture, and the argument stops being about platforms.
If you want the whole architecture in a single pass rather than link by link, the five-layer machine data flow covers the same ground as one continuous walk from PLC to decision owner.
Which link in your chain has cost the most and can name the fewest decisions?
The eight-link data chain, the per-stage review questions, and the portfolio framing of measurement coverage are from Industrial IoT and Data Architecture: From Sensor to Historian to Dashboard (Part 0: The Data Value Chain, sections 0.3–0.4, and section 1.1).
Questions industrial leaders ask about this
What is the industrial data value chain?
An editorial organising model with eight links: process, sensor, signal, acquisition, transport, storage, context, decision. A physical process produces a condition, a sensor turns it into a signal, an acquisition device turns the signal into named and timestamped data, a network moves it, a store keeps it, a context layer gives it meaning, and a consumption layer puts it in front of a decision. It is not a standard and it does not map to one; its job is to make the hand-offs between links visible so a review question can be asked at each.
Why is value created only at the last link?
Because nothing upstream changes an outcome by itself. A reading in a historian has cost engineering hours, hardware, configuration, bandwidth, licence and modelling, and it has changed nothing until a person or a system decides differently because it exists. Every link before the decision is expenditure; the decision is where the expenditure either converts or does not.
What does building the chain backwards mean in practice?
Start at the decision and walk to the physics. Name the recurring decision, the role that makes it, the value or trend that would change it, and roughly what the decision is worth. Only then ask what physical effect carries the earliest useful evidence for that decision, what accuracy the decision actually needs, and where a sensor would have to live to see the effect. Walked the other way — machine first, catalogue open — every question defaults to more, better, everywhere.
What are the two cross-cutting concerns in the chain?
Security, because every link is also an attack surface, and ownership, because every link needs a named owner for its configuration, its failure alarms and its changes. Neither lives at a single stage. A chain diagram with a product logo in every box and no name against any of them has documented a purchase, not an architecture.
Does a completed chain review mean the architecture is right?
No. A completed review shows the questions were asked, not that the answers are correct. The chain is a map of questions that makes missing cost, quality and ownership visible. Selection, installation, configuration and change remain site engineering under applicable law, OEM instruction, controlled procedure and qualified approval.
Industrial IoT and Data Architecture — From Sensor to Historian to Dashboard
The nine-part reference this series draws on: the data value chain, the sensing layer, edge and acquisition, plant networks, historians and time-series storage, context and asset models, dashboards and analytics, security and chain reliability, and the implementation playbook — 40 sections covering signal families, protocol theories, timestamp discipline, compression and retrieval, tag naming, asset models, data contracts, notification engineering, threat modelling, and the pilot-to-wave economics a finance function can audit.