At 09:15 on a Wednesday, the site reliability meeting opens with the monthly bad actor list on screen. Ten equipment tags, ordered by count of unplanned work orders in the last twelve months. The number-one asset is a slurry pump that has produced twenty-three corrective work orders. The engineer who owned last quarter’s action against that pump was reassigned in June, the seal upgrade specified in that action was descoped for cost, and no one in the room can say whether the new seal type would have prevented any of the failures listed. The list moves to the top of next month’s agenda. The pump keeps failing.
This is the shape of most bad actor programmes at the twelve-month mark: an authoritative list produced on a fixed cadence, generating no verifiable retirement of repeat failures. The Pareto observation the programme rests on holds up. Industry practitioners consistently report that roughly ten to twenty per cent of a plant’s critical assets account for the majority of unplanned downtime and unbudgeted maintenance spend at any point in the year. What decides whether the programme moves the number is not the identification, which any competent CMMS report will do. It is the record that carries each bad actor from identification to retirement, and the review that will not close it until the failure signature has stopped.
What the record has to carry
A bad actor register that produces decisions carries more than a rank and a failure count. In IBM Maximo and MAS the working record for each entry is typically held against the asset, with linked work order history and a small set of custom attributes carrying the fields below.
- Asset tag and parent hierarchy, so a ranked pump is visible in the context of the system it disables, not in isolation. Bad actors are frequently pumps and valves whose failure signature belongs to the upstream process condition, not the asset.
- Failure signature at the failure mode, mechanism and cause levels of the ISO 14224 failure hierarchy, drawn from the last twelve months of closed work orders on that asset. A rank derived from mixed corrective and preventive orders, or from unstructured comments, cannot be used to prioritise a remedy.
- Loss basis, stated: unplanned downtime hours, direct maintenance cost, production loss, safety event count, or a weighted composite. The choice of basis is the choice of programme. Changing it mid-year invalidates the trend.
- Consequence class, derived from asset criticality, not overridden by it. A Class C repeat leaker can top a downtime-weighted list; it should not top a safety-weighted one.
- Current remedy hypothesis, one sentence, naming what change is expected to remove the failure signature: seal type change, PM revision, control loop retune, spare part specification change, operator procedure change, capital replacement.
- Owner and by-when, named individuals with a delivery date, not a group inbox or a rolling “under review” status.
- Verification test, defined at the time the remedy is agreed: what will have to be true, on which report, by which date, before this asset comes off the list.
An entry without a verification test is a compliance artefact. It will still be on the list next quarter.
How the ranking is calculated
The choice of ranking algorithm is a design decision, not a data science question. Three approaches survive contact with an operations team.
- Downtime-weighted Pareto. Rank by unplanned downtime hours attributable to the asset over the last twelve months. Simple, defensible, and biases the programme toward availability. Loses safety-critical low-frequency events.
- Cost-weighted Pareto. Rank by direct maintenance cost plus a nominated production loss rate. Biases toward operating margin. Requires trustworthy production loss figures; where those are unavailable, the ranking is illustrative rather than audit-grade.
- Consequence-adjusted Pareto. Downtime or cost ranking multiplied by an asset criticality factor derived from the site’s own criticality framework. This is the ranking that survives being read alongside the risk register.
Whichever calculation is chosen, it is defined in writing, stored in the CMMS as a saved query or KPI, and applied identically each cycle. A quarterly change of algorithm turns the register into a rolling debate rather than a management tool.
Bad actor ranking runs on structured failure data, so the failure code library has to be granular enough that a slurry pump reads differently from a boiler feedwater pump. Where the library collapses everything into “mechanical failure” the ranking becomes a proxy for which asset has the most-used tag on a work order.
The review that turns the register into action
The bad actor list produces retirements when three governance elements are in place, on a fixed cadence, with the same room of people each time.
- Monthly review at reliability engineer level. The top ten by the agreed ranking, plus every entry whose remedy is past its by-when date. The output is a revised remedy, an escalation, or a decision to retire the entry against its verification test.
- Quarterly review at reliability manager and operations manager level. How many entries retired against verification in the last quarter, how many escalated, and the aggregate downtime and cost trend against the twelve-month baseline. This is where remedies that keep failing get reframed as capital or as strategy work rather than as more maintenance.
- Annual review at asset manager level. Whether the ranking basis remains fit, whether new bad actors are appearing at a stable or rising rate, and whether the failure code library is still granular enough to describe what the site is actually seeing. This is where the register connects to the Strategic Asset Management Plan under ISO 55001.
Two operational rules keep the review honest. An entry only leaves the register when its verification test passes on the report specified at the time of the remedy, not on the reliability engineer’s assessment. And a failed remedy is treated as evidence about the diagnosis, not about the engineer; the record is reopened with a revised hypothesis rather than closed as “monitoring”.
Where the analysis stops being useful
Bad actor analysis works on assets with enough failure history to produce a signature. It is a poor lens for genuinely new equipment, for one-off failures on high-criticality assets where a single event dominates the twelve-month record, and for systemic problems that show up across many assets in the same class rather than in one tag. The systemic case is the one that hides most often: a valve series with a persistent packing failure across sixty assets will not appear on a top-ten list by tag because the count is spread thin, but in aggregate it is a larger loss than any of the ranked entries. The control is a second query, run each cycle, that ranks failure signatures across the fleet rather than assets, using the same failure code hierarchy the register depends on.
The register also degrades if root cause analysis is run as an event rather than a programme. Without a functioning RCA function feeding remedy hypotheses back into the record, the list becomes a monthly count parade with no diagnostic depth behind any of the top ten entries.
Two organisational conditions decide whether the register survives its first eighteen months. The reliability function needs a mandate to close a bad actor line on the verification test alone, so that a retirement is a data event rather than a negotiation with the maintenance manager. And accountability sits with one named role at asset director or head of reliability level; distributing it across three functions produces three lists and no retirements.
Closing position
A bad actor programme that retires repeat failures is not a report. It is a controlled register in the CMMS, with named remedy hypotheses, named owners, a verification test on every entry, and a review cycle that will not close a line until the failure signature has actually stopped appearing. The Pareto observation that a small share of assets carries the majority of the loss is durable; whether the programme uses it depends entirely on the discipline of the record. Operators who install this discipline see the twelve-month rolling downtime attributable to the top ten flatten and then fall inside a year. Operators who keep running the list as a count parade will still be talking about the same slurry pump this time next year, with a different engineer’s name against it.