When AI Enters the Operations Room: What Reliability Actually Requires
An original perspective on deploying AI in environments where being wrong is expensive
By Bhavin Gandecha
There is a particular kind of room where I have spent a lot of my working life. It does not always look the same. Sometimes it is a network operations centre with screens covering one wall. Sometimes it is a small meeting room where an incident bridge is being run from a laptop. Sometimes it is just an inbox at three in the morning. What these rooms have in common is that the people inside them are responsible for keeping something working that other people are relying on. When AI shows up in these rooms, and it is starting to show up everywhere, the conversation around it tends to be very different from the conversation in the wider technology press. The hype matters less. The behaviour under pressure matters more.
I want to write about that gap. There is a serious and growing assumption that AI will improve the way operational environments run. In some cases I think this is true. In others I think we are putting tools into the room without thinking carefully enough about what the room actually demands. The purpose of this piece is to set out, in plain language, what reliability actually requires from any AI system that is being asked to support real operations, and what the conditions are that decide whether such a system helps or quietly makes things harder.
The First Question Is Always About Failure
In operations, the most important question about any new tool is not what it does when it works. It is what it does when it is wrong. Human operators learn this the hard way over the course of their careers. They develop a sense of which colleagues to trust under stress, which alerts can be ignored, which suppliers will deliver on time. The judgment is shaped by years of seeing things go wrong and noticing the patterns. When an AI system is introduced into the same environment, that hard won judgment now needs to be extended to a new participant. The team has to learn what the AI is good at, what it is bad at, and most importantly what it looks like when it is bad.
This is a much harder thing than it sounds. A traditional tool, like a monitoring dashboard, fails in obvious ways. A graph stops updating. A threshold is crossed. A red light appears. AI systems often fail in the opposite way. They keep producing confident output even when that output has drifted away from reality. The interface still looks the same. The summary still reads fluently. The recommendation still sounds reasonable. The error is silent, and it is buried in language that mimics the language of correct answers. In an operations room, this is a serious failure mode. Operators trained to look for visible warning signs may miss the quiet ones, and the cost of missing them can be significant.
Reliability Is Not the Same as Accuracy
It is tempting to evaluate an AI system by its accuracy rate. If a model is correct ninety five percent of the time on a benchmark, that sounds impressive, and in many contexts it is. But operations work is not a benchmark. The same accuracy rate can produce wildly different outcomes depending on what the system is being asked to do, how often it is being asked to do it, and what the consequences of the remaining five percent look like.
Consider a network operations setting where an AI assistant is asked to suggest the most likely cause of an outage. Ninety five percent accuracy on a known set of historical incidents might be acceptable for a triage tool. The same accuracy rate becomes alarming if the tool is being used to decide whether to fail traffic over to a backup region, because the five percent of wrong answers may include the ones that matter most. The cases where an outage looks like something familiar but is actually something new are the cases where humans need to stay sharp, and they are also the cases where the AI is most likely to confidently suggest the wrong thing.
Reliability, in the operational sense, has to account for this asymmetry. It is not about being right most of the time. It is about being wrong in ways that can be caught, and about not being wrong in ways that propagate.
The Role of Degraded Operation
Every well designed system has a concept of degraded operation. A network can lose half its capacity and still carry priority traffic. A hospital can lose its primary record system and still treat patients with paper for a few hours. A financial platform can fall back to a slower path while a faster one is being repaired. The degraded state is not a failure. It is a planned, rehearsed, and survivable mode that buys time for recovery.
AI systems in operational settings need their own version of this. What does the operations room look like when the AI assistant is down. What does it look like when the AI is up but its outputs are no longer trustworthy because the model has drifted, or because the underlying data has changed, or because an upstream provider has changed something quietly. The honest answer in many teams today is that nobody has thought this through. The AI is treated as either present or absent. There is no middle state. There is no defined moment at which an operator should stop trusting it and switch to a manual approach. There is no rehearsal of what that switch looks like under stress.
This is a serious gap, and it is one of the easiest to close. A team that has written down, even briefly, what their degraded AI state looks like, who declares it, and what changes when it is declared, will behave very differently in a real incident from a team that has not. The work is mostly process work, not technical work, and it is one of the highest value things an operations leader can do this year.
Data Conditions Decide Outcomes
It is often said that AI is only as good as its data, but in operational settings the point is sharper than that. The data conditions in which the AI was trained are usually not the data conditions in which it is being asked to operate. Training data tends to be cleaner, more complete, more representative of normal conditions, and labelled by people who had time to be careful. Operational data is messy, partial, sometimes wrong at source, and often arriving under conditions where nobody has time to clean it. The gap between these two states is where a lot of the disappointment with AI deployments lives.
Closing this gap is harder than it looks. It requires honest documentation of what data the model expects, what data the operational environment actually produces, and what compensating controls bridge the difference. It also requires somebody to own the question of when the gap has become too wide. A model that was trained on traffic patterns from one year may not be sound for the next year if the network has changed shape. A model trained on incident data from one organisation may not transfer well to another. These transitions need to be tracked, and they need to be tracked by name, by date, and by responsible owner. That is not a research activity. That is an operations discipline.
Human Authority Has to Stay Clear
There is a quieter risk in introducing AI into operations rooms that is rarely discussed. It is the gradual erosion of human authority over decisions that ought to remain human. When a system produces a recommendation, and that recommendation is correct most of the time, the people in the room start to lean on it. Over time, the lean becomes a habit. Over more time, the habit becomes the default. By then, the original judgement that the humans were supposed to bring has quietly faded, and the AI is no longer a tool. It is the decision maker, with humans serving as approvers who no longer have the depth to question it.
This is not an argument against using AI in operations. It is an argument for being deliberate about which decisions remain human, and for protecting the human capability needed to make those decisions well. Operators need to keep practising the work that the AI is now doing for them most of the time, in the same way that pilots still hand fly the aircraft during training even though most of a real flight is automated. If we lose that practice, we lose the ability to take over when the automation is wrong, and we will only discover the loss in the worst possible moment.
What Good Looks Like
A well integrated AI capability in an operational environment looks calm, not impressive. It does fewer things than the marketing material suggests, and it does them with clear boundaries. The operators using it can describe in plain language what the tool is good at, what it is not good at, and what to do when it is uncertain. The team has a written degraded state. The data conditions are documented. The decision rights between humans and the model are clear, and they are revisited regularly. The model is updated on a known cadence, and updates do not arrive silently in the middle of an incident.
Most importantly, the people in the room still know how to do the job without the AI. They might be slower without it. They might be less consistent. But they are not lost, and they can recognise when the AI itself is the problem rather than the solution. That last capability is the single most valuable thing an operations team can preserve as AI tools spread, and it should be defended deliberately, not assumed.
Closing Thought
The conversation about AI is dominated by what is possible. The conversation about operations has always been dominated by what is reliable. Both conversations are right, but they are pointed in different directions, and the work of bringing them together is being done quietly, in unglamorous settings, by people who have not been given the language to describe what they are doing. My hope is that more of that language finds its way into the public conversation, because the decisions being made now about how AI enters operational environments will shape the reliability of important systems for the next decade and beyond.
The framing that matters is not whether the AI is clever. It is whether the system that contains the AI, the people, the data, the processes, the boundaries, can be trusted to hold up when it counts. That is what reliability actually requires, and it is the conversation that I think is most worth having now.