Home / Impact / Representative engagement
Representative engagement Mining & Processing · Reliability-Centred Maintenance (RCM2)

Rebuilding maintenance strategy around failure modes

How a processing operation moves from a reactive, calendar-driven maintenance plan to a defensible strategy its own engineers own and sustain — and why most such programmes fail before they get there.

5 days
In-house cohort
RCM2
Methodology · SAE JA1011
Mixed
Operations + maintenance
Sector
Mining & Processing
Asset class
Continuous processing line
Programme
RCM Level 2
Format
In-house, on client assets
Standards
SAE JA1011 · ISO 55001

The situation we are usually called into

A processing operation runs a continuous line whose availability determines the entire production plan. Maintenance is competent and hard-working — and almost entirely reactive.

Interventions are calendar-based, inherited from OEM manuals and adjusted over the years by whoever won the argument. Failure data exists in the CMMS but is coded inconsistently, so nobody can say which failure modes actually cause the losses. Each shutdown is investigated as an isolated event.

This is not unusual. It is the modal state of the industry.

What the benchmarks say
  • Top-performing plants target under 20% reactive maintenance; most facilities are the inverse of that target.
  • Reactive-dominant operations are associated with materially higher defect rates and maintenance-related delays.
  • Roughly 10% or less of industrial equipment ever actually wears out — meaning most mechanical failures are, in principle, avoidable.
  • Maintenance can absorb 20–60% of operational expenditure, depending on industry and asset type.
Sources: U.S. Department of Energy benchmarks; McKinsey; Plant Engineering / industry surveys.

The real problem is not maintenance. It is decision-making.

There is no shared, defensible basis for deciding what to maintain, when, and why. Three consequences follow, and they compound.

The task list only grows. Nobody can justify deleting a preventive task, because no analysis exists that would tell them it is worthless. So the plan accumulates work that prevents nothing — while the failures that actually cause losses remain unaddressed. High schedule compliance and high failure rates coexist comfortably, because compliance measures whether the plan was followed, not whether the plan was right.

Critical knowledge is personal, not institutional. Failure modes on the crushing and screening circuits are managed by the experience of a handful of people. That experience leaves with them.

Hidden failures are never counted. Protective devices, standby equipment, trip systems and relief paths fail silently. Nothing announces it. Until a second event demands the failed function, the organisation is carrying risk it has neither quantified nor decided to accept. No amount of condition monitoring on running equipment reveals it.

“The analysis is almost never the hard part. Getting an organisation to be obliged to act on it is.”

How we run the programme

RCM2 is delivered in-house to a deliberately mixed cohort — operators and maintainers together, because the people who run an asset hold knowledge the people who maintain it do not. The analysis is performed on the client’s own critical assets during the engagement, not on a textbook case.

Define the operating context

What the asset is required to do, against the performance standards the business actually needs — throughput, availability, product specification, environmental limits. Everything downstream depends on getting this right, and it is routinely skipped.

Identify functional failures

Including partial failures, where the asset still runs but outside acceptable performance. These are the losses that never appear in a downtime report.

Build the FMECA

Failure modes and effects for each functional failure, with consequences assessed across safety, environmental, operational and economic categories — using the organisation’s own risk framework, so results are comparable across the plant.

Select failure-management policies

By consequence, using RCM2 decision logic: condition-based tasks where a P-F interval can be established; scheduled restoration or discard where age-related failure is demonstrable; failure-finding tasks for hidden failures; redesign where no task is both technically appropriate and worth doing.

Separate hidden from evident failures

And derive failure-finding intervals from the required availability of each protected function. This is arithmetic, not judgement — and it is the work most sites have never done.

What makes it survive the next budget cycle

The analysis is not the deliverable. Programmes that fail have good analyses sitting in folders. Three things are built alongside it:

Decision rights
Named approvers for interval changes, task deletions and new failure-finding tasks — each with a deadline. Recommendations that miss it escalate rather than expire.
A path into the system
Every approved recommendation enters the CMMS as a dated, audited change. If it does not reach the system, it does not exist.
Failure-mode coding
Work orders re-coded so future failures attribute to the modes the analysis identified — closing the loop that previously prevented learning.

What the client is left with

A maintenance strategy that can be defended, asset by asset, to an auditor or a board: this task exists because this failure mode has this consequence, and this policy is technically appropriate and worth doing.

More durably, the capability stays in-house. The client’s own engineers facilitate analyses on new assets and propagate findings across circuits. That is the mechanism by which reliability compounds instead of resetting with each change of leadership.

Programme
Rebuilding maintenance strategy around failure modes at a diamond processing operation
View programme

Could this work for your operation?

Discuss an in-house programme