MTBF: The Most Misread Number in Engineering
A drive datasheet quoting 1–2.5 million hours appears to promise 114 to 285 years of life. The reported fleet reality is an annualised failure rate of a percent or two. Both numbers are true, because MTBF is a population rate restated in hours — never a lifespan.
A hard-drive datasheet quotes an MTBF somewhere between 1 and 2.5 million hours. Divide by the 8,766 hours in an average year and the drive apparently lasts between 114 and 285 years.
Nobody believes that when it is said out loud. The same arithmetic runs unchallenged on plant equipment every week.
Backblaze, the cloud-storage company, has published quarterly failure statistics for its own disk fleet — hundreds of thousands of drives — since 2013. Its reported fleet reality sits in the low single-digit percent: annualised failure rates broadly in the 1–2% band across recent quarterly reports, moving with drive model, age mix and quarter, and climbing in the older cohorts as they age.
Both numbers are true. Neither is a specification for anyone else’s fleet.
The MTBF describes a failure rate across a large population during useful life. It says almost nothing about how long one unit lives. An engineer who confuses the two schedules overhauls that make nothing better, and writes business cases that cannot survive contact with a CMMS.
The working set, and the trap in each
Six quantities do nearly all the work in industrial practice, and each carries a standard misreading.
| Quantity | What it actually is | The misreading |
|---|---|---|
| Failure rate λ(t) | failures per unit of operating time in a population, at age t | assuming it is constant — that is a special case, not a law |
| MTBF | mean operating time between failures for repairable items: a population rate restated in hours | reading it as service life |
| MTTF | the same idea for items replaced rather than repaired — what a disk datasheet is really describing | using the two words interchangeably |
| MTTR | mean active restoration time: the hands-on work, once people, parts and permits are present | quoting wrench time and calling it downtime |
| Mean downtime (MDT) | the whole outage: detect → diagnose → wait for parts, permits and people → repair → return to service | leaving it out of the availability figure |
| Availability | inherent: MTBF ÷ (MTBF + MTTR). Operational: mean uptime ÷ (mean uptime + MDT) | mistaking either for reliability |
Both availability figures are steady-state arithmetic. They assume rates stable enough to average, and they count only the downtime you put in the denominator — which makes the choice of denominator the whole argument. Inherent availability flatters a plant: it counts the spanner hours and ignores the night spent waiting for a courier. Operational availability is the number the plant actually lives in.
Quote the first for a good-looking KPI. Quote the second when you want the one a predictive programme can move.
And note what the identity says once the terms are honest. There are exactly two levers: fail less often, or restore faster. Early warning can pull both — heading off some functional failures where the mechanism is detectable early enough and someone acts on the warning, and converting others from unplanned, with its diagnosing and sourcing and waiting, into planned, with parts staged and a window booked. That double effect is the mechanism behind the programme-level savings ranges that circulate in this field — ranges compiled from self-selected programmes, with unstandardised accounting and no counterfactual behind them. Whether anything of that size transfers to your plant is a separate question, and no published range answers it for you.
A number without a population is not a number
Reliability arithmetic only means something across a population — a fleet of identical motors, a pump model on one duty class — and inside an operating context. The same pump model has a different λ on abrasive slurry than on clean water, and averaging across both produces a number that describes neither.
Two consequences bite in practice.
Small plants have small numbers. Two failures in three years is not a failure rate. It is a scrap of evidence with error bars wide enough to drive a truck through. This is why criticality workshops score likelihood by judgement and say so, why a failure-mode library pools modes across comparable assets, and why industry-scale databases such as OREDA exist at all.
Most of your data is censored. The motors that have not failed yet are evidence too — evidence of survival. Ignoring them and counting only failures biases every estimate pessimistic. Survival models are built to use censored records properly, and cleaning the CMMS is what makes install dates and run hours trustworthy enough to try.
Where the hours actually go
At Meridian — the illustrative teaching plant used throughout this series, invented rather than observed — the screw compressor K-201 had two recorded restorations in its history. Each ran about fourteen hours. Under four of those were hands-on repair. The other ten went to noticing, diagnosing, and waiting for parts.
Two events cannot estimate an MTTR any more than they can estimate a rate, and the discipline applies to a teaching example as much as to a real one. But two events can show where the hours went, and that is a description rather than an estimate.
It is enough to reframe the argument. If ten of every fourteen hours sit outside the repair, the business case was never mostly about wrenching. A warning that arrives early enough to stage the job compresses the waiting. How much of those ten hours a real warning recovers depends on how much lead time the failure mode gives you and whether the planning system uses it — two things a spreadsheet cannot assume on your behalf.
The practical move is unglamorous: split downtime into its phases in the CMMS — detect, diagnose, logistics, repair — because a predictive programme shows up as collapsed detect-and-logistics time, and a KPI that only records total downtime cannot see it happen.
MTBF alone cannot schedule an overhaul
That decision needs the failure pattern, and a test of whether the proposed task is applicable to the mode and worth doing for the consequence. MTBF alone has justified a great deal of intrusive preventive maintenance that the pattern never supported.
The evidence for treating that as a real risk is old and repeatedly re-examined. United Airlines’ component data, analysed by Nowlan and Heap and published in 1978, found that only about 11% of the component types studied showed age-related failure; 89% did not, and the single largest share, 68%, was an infant-mortality pattern in which failure risk is highest immediately after installation or overhaul. Replications followed — Bromberg’s Swedish study in 1973, the U.S. Navy’s Maintenance Steering Group analysis of naval aircraft in 1982, the Navy’s submarine programme in 2001. As those results are compiled in the RCM literature, the age-related share lands somewhere between roughly one component type in twelve and a little under one in three.
Treat the exact splits as indicative: the compilations differ in detail, not every primary study is easy to obtain, and none of these are industrial plant. What is durable is the direction, and the question it forces with your own PM schedule open — does this failure mode have a wear-out zone to aim a calendar at, and how many healthy machines am I opening to find out?
The corresponding thing to hunt in your own records is a cluster of failures in the weeks after planned work. It is one of the most actionable findings a work-order history can hand you, because the countermeasures — alignment, torque, cleanliness, commissioning checks — cost almost no capital. It also has innocent explanations that have to be ruled out first. The PM may have been triggered by a developing fault, and return-to-service may simply be the only time anyone is watching closely.
An exercise that costs an afternoon
Pick one asset class where you have enough units to argue about — the forty identical motors, one pump model on one duty. Not one machine.
Export every failure record for that population, then do three things with it.
Write the population and the context beside the number. “MTBF of this pump model, on this duty,” never a bare figure. If you cannot state both, you have found the reason the number has been travelling further than it should.
Count the survivors. Pull install dates and run hours for the units that have not failed. That column is usually missing, and its absence is what makes every estimate in the file pessimistic.
Split the two or three longest outages into phases. Detect, diagnose, logistics, repair, from work-order timestamps and the technician’s notes. The awkward one belongs in the sample.
What usually happens is that the arithmetic refuses to support a rate, which is a finding rather than a failure — say it out loud, use a pooled library value where one exists, and flag the judgement for replacement as history accumulates. And the phase split usually shows most of the hours sitting outside the repair, which quietly rewrites what the next improvement should be aimed at.
Reliability numbers mean something only with a population, a context and a shape attached. MTBF is a rate restated in hours, never a lifespan.
What population and operating context sits behind the MTBF on your dashboard — and who wrote it down?
The reliability working set, the Backblaze restatement, and the population-and-censoring disciplines are from Predictive Maintenance: Practitioner Reference Frameworks and Planning Guide (Part 3: Failure Modes, Reliability, and Degradation).
Questions industrial leaders ask about this
What does MTBF actually mean?
MTBF is the mean operating time between failures for repairable items — a failure rate across a large population during useful life, restated in hours. It is not a service life for one unit. Manufacturers quote drive MTBF figures of 1 to 2.5 million hours; divided by the 8,766 hours in an average year that implies 114 to 285 years, while Backblaze's reported annualised failure rates for its own fleet sit broadly in the 1–2% band. Both numbers are true and neither is a lifespan.
What is the difference between MTBF and MTTF?
MTBF applies to items that are repaired and returned to service; MTTF applies to items that are replaced rather than repaired, which is what a disk datasheet is really describing. Using the two words interchangeably loses the repairable-versus-replaceable distinction, and with it the reason a fleet statistic cannot be read as the life of the unit in front of you.
Is MTTR the same as downtime?
No. MTTR is mean active restoration time — the hands-on work, once the right people, parts and permits are present. Mean downtime is the whole outage: detect, diagnose, wait for parts, permits and people, repair, return to service. Quoting wrench time and calling it downtime hides exactly the hours a predictive programme acts on, which is why downtime is worth splitting into detect, diagnose, logistics and repair phases in the CMMS.
How many failures do you need before MTBF means anything?
More than a plant usually has for one machine. Two failures in three years is not a failure rate; it is a scrap of evidence with very wide error bars, and two events cannot estimate a mean restoration time either. Reliability arithmetic needs a population — a fleet of identical motors, or a pump model on one duty class — and an operating context, because the same pump model behaves differently on abrasive slurry and on clean water.
Can MTBF decide when to overhaul a machine?
Not on its own. That decision needs the failure pattern and a test of whether the task is applicable to the mode and worth doing for the consequence. In the aviation dataset Nowlan and Heap analysed, about 11% of component types showed age-related failure and 89% did not. Those are that fleet's proportions, not yours, but the question they force is general: does this failure mode have a wear-out zone to aim a calendar at?
Predictive Maintenance — Practitioner Reference Frameworks and Planning Guide
The twelve-part reference this series draws on: foundations and the value case, asset criticality and strategy, failure modes and degradation, the monitoring technologies, asset-class playbooks, sensors and IIoT architecture, data foundations, signal processing, analytics and prediction models, alerts and diagnosis, work management and CMMS integration, and pilot execution through rollout and governance — 126 sections with 46 technical figures.