backicon

From Fault to Outcome: How the Autonomous Maintenance Workflow Works

Published on :

August 18, 2026

by

Anisha Bhattacharjee

view

Views...


What does "autonomous" actually mean in autonomous maintenance? At Xempla, the boundary is simple: everything up to the point where someone needs to physically pick up a wrench and go to the asset should happen with minimal human supervision. That raises the more useful question: what actually happens between the moment something changes on an asset and the moment it becomes a resolved outcome, and who does that maintenance event pass through along the way?

At Xempla, this is what autonomous maintenance is actually built to do, and the way we've structured it is through four stages: Discover, Investigate, Implement, and Verify, what we call the DIIV cycle, a framework we've written about separately. But the story here isn't the framework itself, it's what that framework looks like in motion: an operation involving different people, different systems, and different levels of responsibility, each relating to the same issue differently, at a different point in its life. A supervisor and a facility manager are, in a very real sense, looking at the same event, just from entirely different altitudes.


It starts with an asset, not a work order

An asset changes before anyone necessarily knows something is wrong. Omi, Xempla's reliability agent, continuously monitors assets for exactly this kind of deviation, watching for the kind of drift that doesn't announce itself the way an outright failure does. The moment Omi detects one, the event has effectively entered the operation, well before any person is aware of it.

Omi doesn't simply flag that something changed. It assesses the evidence and triages the event into one of three paths: work order generation when the evidence supports immediate action, investigation when the situation needs closer engineering review before a decision can be made, or observation when the deviation isn't yet significant enough to act on but is worth tracking. Along the way, it also does the quieter work of filtering out false alarms, sensor noise, seasonal variance, one-off blips, so the same anomaly doesn't keep resurfacing as if it were a new problem each time.

There are also cases where the evidence doesn't clearly support one path over another, where Omi isn't confident enough to make the call itself. When that happens, it doesn't just pass the case along. It flags the two or three specific things worth checking, so the engineer starts with a shortlist instead of a blank investigation.

When unsure, Omi narrows the investigation by clearly identifying what needs to be checked. 


The case reaches engineering

From here, the case reaches the engineering layer, which Xempla calls the ROC, the Remote Operations Center. Depending on how a customer has set things up, this can be Xempla's own ROC team acting as that layer, or the customer's own engineering team using the platform to do the same job. Either way, what changes is what the engineer is starting from.

The engineer no longer begins with a raw signal and has to build the case themselves, piecing together asset history, prior incidents, and context from scratch. The case arrives already assembled, with the evidence, the reasoning, and a confidence score attached. For issues Omi has flagged for a work order, the engineer remains the checkpoint before anything is dispatched. For issues flagged for investigation, the engineer goes deeper, using the context already gathered to reach a decision the system alone didn't have enough certainty to make on its own.

The engineer can accept the recommendation, add to it, or override it. That input feeds back into how the system handles similar cases going forward, across comparable assets, not just the one in question. This is also where our previous piece on who actually gets notified once maintenance is autonomous connects in, because at this point the case has stopped being a system output and become something a specific person is responsible for.


Decision and execution are not the same step

A decision tells you what needs to happen. Getting that decision scheduled, prioritised, and slotted into everything else already in motion at a facility is a separate piece of work, and that's the job Nira, Xempla's planning and scheduling agent, does. Nira checks the new work against existing PPM schedules, open work orders, and current priorities, and recalibrates the plan so it lands in the right place rather than competing for attention on its own.


The work reaches the technician

The decision has now been translated into an actionable piece of work, and it reaches the person who executes it physically. For hard FM, this happens through the Xempla mobile app, used on the ground by technicians, so the person arriving at the asset already has the context they need, what the issue is, what's been recommended, what to check, instead of working it out after they've arrived. This is the point where physical intervention has to happen, and no amount of upstream automation changes that. Someone still has to be there, hands on the asset.

The technician gets guided tasks and can capture evidence at the asset. 

That technician's work sits inside a larger shift that they're not responsible for managing alone. Supervisors oversee both hard FM and soft FM together across their shift, work orders, PPM, tickets, checklists, all prioritised and visible in one place, rather than pieced together manually from separate systems and separate conversations. This is where hard and soft FM genuinely connect. Autonomous maintenance, in the sense we've defined it, is most directly applicable to hard FM, where faults appear as measurable, trackable deviations on physical assets. Soft FM follows a different operational path, governed by tickets, checklists, and quality checks rather than continuous asset monitoring, but it sits within the same shift, managed by the same person, contributing to the same overall picture of whether a facility is running the way it should.


From technician to portfolio, the question keeps changing

Follow the same event upward and the question being asked about it changes at every level. The technician asks what needs to be done at this asset. The supervisor asks what needs to happen across this shift, given everything else on their plate. The facility manager, a level above, isn't asking about this one event specifically anymore, they're asking whether the facility is performing, whether SLAs are on track, and whether what just happened is a one-off or part of something recurring. 

At the facility level, the focus shifts from individual tasks to overall performance. 

At the portfolio level, the question becomes broader still: what patterns are emerging across facilities, and what do they say about the operation as a whole. Each role is looking at the same underlying event, but with a different question in mind.


The event doesn't disappear when the work order closes

Closing a work order isn't the same as resolving the problem, and treating it that way is one of the more common gaps in traditional maintenance operations.

For hard FM, the primary verification comes from continued monitoring. Omi returns to the asset after the work order closes and checks whether it actually recovered to the condition it should be in. If it has, the loop closes cleanly. If the deviation is still there, the case reopens and goes back into the workflow rather than being treated as resolved simply because the work order was closed. Xempla's Assurance Score adds another layer to this, reading the operational health of a facility and its assets on an ongoing basis, independent of any single event, rather than relying on one closed work order as proof that things are fine. Separately, the workflow also preserves better evidence of what actually happened at the asset: the mobile app requires real-time documentation, including photos, before a work order can even be submitted, so the record doesn't depend on someone's memory of it after the fact.

For soft FM, the equivalent check works differently, since there's no physical asset behaviour to monitor afterward. Instead, the work logged is reviewed against defined quality requirements, and anything that falls short gets flagged rather than assumed to be fine simply because a ticket was closed.


From one event to portfolio learning

One fault, on its own, cannot tell you whether a portfolio is healthy. But hundreds or thousands of verified events, accumulated over time, start to show something a single event never could: where problems recur, which assets keep failing in the same way, and where performance is genuinely shifting across sites rather than just fluctuating.

What makes this possible is that the event is never actually lost as it moves upward, it just changes shape. The same deviation that began on one asset can be seen, at the same time, as a maintenance task by the technician, as one item in a shift by the supervisor, as part of facility performance by the facility manager, and, over time, as one contributor to a portfolio-level pattern. It's the same event throughout. Only the altitude, and the question being asked of it, changes.

The goal of autonomous maintenance was never to remove people from this journey. It is to make sure every person who touches it, engineer, technician, supervisor, or facility manager, enters with the context, priority, and information they need, while the system carries the work forward between them and verifies the result at the end.

A signal starts the process. It becomes a decision, then a plan, then physical work, then a verified result, and finally, one data point in how the portfolio is understood. That's how the workflow actually works.


FAQs

What is autonomous maintenance?

Autonomous maintenance at Xempla means everything up to the point where someone needs to physically intervene at the asset happens with minimal human supervision. The workflow forms a continuous loop, moving from detection and investigation to decision, execution, and verification, with the outcome feeding back into ongoing monitoring.

How does autonomous maintenance handle issues when the system is uncertain?

When the evidence is not sufficient to make a reliable decision, handover analysis gives the engineer the relevant context, confidence level, and specific checks needed to narrow the investigation.

What role do engineers play in autonomous maintenance?

Engineers act as a decision checkpoint for cases that require human judgment. They can accept, modify, or override system recommendations, with their input helping improve how similar cases are handled in the future.

How does autonomous maintenance support technicians?

Technicians receive actionable work with the relevant context and guided tasks through the mobile workflow. They need to capture measurements, photos, and other required evidence while working at the asset.

How does autonomous maintenance verify that a maintenance issue is actually resolved?

After a work order closes, the asset continues to be monitored to check whether it has returned to its expected condition. If the deviation persists, the issue can re-enter the workflow rather than being considered resolved simply because the work order was closed.

What is Assurance Score in autonomous maintenance?

Assurance Score provides an ongoing view of the operational health of a facility and its assets. It looks beyond individual work orders to assess whether assets and operations are continuing to perform within expected conditions.

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Paragraph

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Bold text

Emphasis

Superscript

Subscript