backicon

Exception-Based Maintenance: Why the Shift Improves Outcomes

Published on :

July 14, 2026

by

Anisha Bhattacharjee

Overnight alerts have piled up on the BMS. Three work orders from yesterday are still open. The CMMS shows a maintenance task due today that may or may not have already been done. Before any real work can begin, the engineer has to piece together what is actually happening. Which alerts matter. Which are duplicates. Which asset has a history that explains the anomaly, and which one is genuinely new. By the time that picture is clear, technicians are already waiting for instructions, new alarms have started coming in, and yesterday's problems still are not fully understood.

This is not unusual. It is a regular morning, and it repeats the same way, every day, on every shift, across every building in the portfolio.

The reason this review takes so long is not a lack of effort. It is that the engineer has effectively become the integration layer between every system in the building. Every alert, work order, maintenance record, and sensor reading eventually converges in one place: the engineer's head. As buildings get more connected and generate more data, this mental integration work grows faster than the team responsible for doing it.


Outcomes Are the New Bar 

Buildings are no longer judged only on whether they were maintained. They are increasingly treated as assets that directly influence business performance. Reliability, energy performance, occupant experience, compliance, and cost efficiency are increasingly evaluated together. Contracts are moving, gradually, from paying for activity to measuring outcomes, and even where a contract has not changed to reflect that yet, the expectation from ownership already has. In this environment, the hours spent every morning reconciling disconnected information are not simply administrative effort. They directly reduce the time available to actually improve the outcomes facilities teams are now being measured on.

The instinctive response to this pressure is to reach for more technology. BMS, CMMS, and IoT sensors have already given facilities teams far more visibility than they had a decade ago, and that visibility is genuinely valuable. But visibility is not the same as a decision. More sensors tend to mean more alerts, and more alerts still need someone to interpret them, connect them to what came before, and decide what to do next. The tools are necessary. The gap between data and decision is what remains.


What's Left Once Review Is Already Done 

Now picture the same morning, but most of it has already been handled before the engineer opens the queue. The overnight alert had already been identified as self-correcting, so it never entered the engineer's queue. A recurring sensor issue was flagged as a known fault, not a new one. What is left is the one thing that genuinely needs a person's judgment, arriving with the context already worked out.

Imagine if engineers only needed to look at the situations that genuinely required human judgment, and everything else was already handled by the time they sat down. This is the idea behind exception-based maintenance: routine cases are resolved without needing an engineer's time, and engineers step in only on the exceptions, the cases that genuinely cannot be resolved without their judgment. This is the foundation of what Xempla calls a System of Decisions: a layer that sits above the existing FM stack, pulling in data from BMS, CMMS, and IoT systems already in place, and turning that data into a small number of clear decisions instead of a growing list of things to individually check.

A decision here is not the same as an alert or a work order. An alert simply flags that something happened. A decision is the conclusion that connects history, related equipment behavior, and probable cause into a recommended action, along with the confidence to act on it.


How This Plays Out Across a Real Day

The queue from the opening doesn't disappear, it just arrives differently. By the time the engineer sits down, the alerts that would once have taken the morning to trace back through history and cross-check against related equipment have already been worked through. What is left is fewer in number, and each one already carries the reasoning behind it, so the engineer's first task is no longer to reconstruct context but to weigh a decision. This is what Xempla's Reliability AI Agent, OMI, does in the background: it triages incoming signals, assembles asset history, related equipment behavior, and probable cause, and arrives at a recommended decision with a confidence score attached. OMI does not take action on its own unless it has been configured to. What it does instead is bring everything, the signal, the context, the reasoning, and the recommendation, into one place, so the engineer sees the full picture at once rather than having to assemble it themselves.

In one of the portfolios Xempla manages, a site with a single 10 kWp inverter was generating consistently between 6.5 and 7 kWp, a persistent and significant drop against its rated capacity. OMI's assessment showed a persistent, significant drop in both generation and inverter input currents, over 90 percent below expected levels, along with voltage instability alongside the drop.

With a confidence score of 0.95, OMI found no other factors that accounted for the scale or suddenness of the deviation, and recommended a corrective work order, with the full justification attached: the specific signals it checked and why the pattern pointed to a genuine operational fault rather than something that would resolve on its own.

OMI's alert on a PV generation drop, showing a 0.95 confidence score, the underlying justification, and a recommendation to generate a corrective work order.

This is the difference between a system that raises alerts and one that reaches a recommendation with the decision already reasoned through. The engineer wasn't left to notice the dip, trace it back through two data sources, and work out on their own whether it warranted a work order. That groundwork, and the reasoning behind it, was already done by the time it reached them, leaving them to make the final call with everything they needed already in front of them.

This is one case, but the same triaging and reasoning runs by default on every signal, across every asset in a portfolio. Across Xempla deployments, this adds up at scale. Human supervisory effort has dropped by 60 to 65 percent, not because less work is happening, but because far less of it needs a person to sit through it first. Between 40 and 45 percent of work orders now run from start to finish without any human review, and fewer than 10 percent of cases ever reach a person for triage at all. The exceptions, in other words, have genuinely become the exception.

The same shift shows up in planning. A maintenance schedule set weeks in advance does not know what has already happened on the ground, so an engineer might still work through a PPM task for something that was effectively resolved days earlier. Nira, Xempla's planning agent, keeps the schedule aligned with what each asset actually needs on an ongoing basis, so the plan reflects reality instead of a fixed date on a calendar.

Execution follows the same idea. Work orders for PPM and fault resolution are filled through Xempla's mobile app, with evidence and real-time context required at the point of work rather than reconstructed later from memory. Once submitted, entries go through an AI-based quality check, and anything inconsistent is flagged for review. Supervisors and engineers get a live view of this work through the Cockpit on the same mobile app, without needing to chase updates manually.

None of these tasks, triaging alerts, adjusting the plan, collecting evidence, checking quality, is particularly difficult on its own. The challenge is that they happen continuously, across hundreds of assets, every single day. That is where the strain quietly builds, long before it shows up as a missed outcome.


Trust Is Built Case by Case 

No organisation reaches this point overnight, and it shouldn't expect to. What builds trust in a system like this is not a single good outcome, but the same outcome holding up, case after case, over time.

AI should not replace engineering judgment. It should protect it. By taking ownership of routine review, investigation, and coordination, a System of Decisions lets engineers spend more of their time where they create the most value: making the decisions where experience, context, and accountability matter most.


FAQs

What is exception-based maintenance?

Exception-based maintenance is an approach where routine cases are resolved automatically, and engineers only step in on the exceptions, the cases that genuinely require human judgment. Instead of reviewing every alert manually, engineers see a smaller number of decisions that already carry context and reasoning.

How is exception-based maintenance different from traditional FM alerting?

Traditional alerting tools flag anomalies but leave the engineer to trace back history, cross-check related equipment, and decide what matters. Exception-based maintenance adds a decision layer on top of that data, so the triage, investigation, and reasoning happen before the alert ever reaches a person.

Does exception-based maintenance replace the BMS or CMMS?

No. The BMS, CMMS, and IoT sensors an FM team already uses continue to do what they do. A System of Decisions sits above these systems, pulling in the data they generate and turning it into a smaller number of clear, reasoned decisions instead of a growing list of individual alerts to check.

How much manual review does exception-based maintenance actually reduce?

Across Xempla deployments, human supervisory effort has dropped by 60 to 65%, largely because most cases no longer need a person to sit through them first. Between 40 and 45% of work orders now run from start to finish without human review, and fewer than 10% of cases ever reach a person for triage.

Does an AI agent take action on its own in exception-based maintenance?

Not unless configured to. In Xempla's model, the AI agent (OMI) triages signals, assembles asset history and probable cause, and arrives at a recommended decision with a confidence score attached, but the engineer makes the final call.

Why does exception-based maintenance matter for facilities management outcomes?

As buildings are increasingly judged on reliability, energy performance, and cost efficiency rather than just maintenance activity, the time engineers spend reconciling disconnected alerts directly reduces the time available to improve those outcomes. Exception-based maintenance frees that time by handling routine review automatically.

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Paragraph

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Bold text

Emphasis

Superscript

Subscript