Nine in ten major IT failures involve human error. Almost none involve carelessness — and the fix is more often the procedure than the person.
When something fails, most organisations fix the person. The data says the procedure is just as likely to be the problem.
A change goes through every stage. Requirement, design, testing, approval, release. Every step is signed. It goes to production and works. A month later it fails. The review concludes that the procedures need updating.
That sequence will be familiar to anyone who has run a change process. Nobody cut a corner. Nobody was careless. Every gate was passed by someone competent doing what the process asked. The outcome was still a failure, arriving late enough that the connection to the change was not obvious at first.
The finding is almost always the same, and almost always vague: update the procedures. Which procedures, updated how, and why they were wrong in the first place, tends not to survive into the action log.
There is now reasonable data on how often this happens and what sits behind it. Start with the encouraging part, because it is real.
Failures are becoming less common. In the Uptime Institute's Global Data Center Survey 2025, half of operators said they had experienced an impactful outage in the previous three years. That is the lowest figure recorded since 2020, when the same question returned 74 per cent.
Uptime Institute Global Data Center Survey 2025. Sample sizes by year: 494, 730, 730, 781, 768, 754. Five consecutive years of improvement.
Investment in equipment, monitoring and process has worked. That is the story of the last five years, and it deserves saying before the rest.
Now the part worth sitting with. Among operators who had an impactful outage in the past three years, 87 per cent said it could have been avoided with better management, processes or configuration. That is up seven percentage points on the previous year. It is a small base — 98 respondents — and it is their own judgement rather than an independent assessment.
Nearly nine in ten failures were inside somebody's control.
Read that the pessimistic way and it is damning. Read it the useful way and it is the most encouraging number in the report. Not weather, not fate, not a supplier. Management, process, configuration.
Which raises the obvious question. If these failures were preventable, and the people involved were competent and following the process, what exactly is failing?
In the Uptime Institute Data Center Resiliency Survey 2026, 92 per cent of respondents said human error contributed at least something to their most recent significant outage (n=220). Thirty-one per cent called it a major contributor, 30 per cent moderate, 31 per cent minor. Only 8 per cent said it contributed nothing.
That number gets repeated a lot, usually badly. So note what Uptime itself says about it. They do not count human error as a root cause. They treat it as a contributing factor, because attribution is genuinely hard. A failed update might be a software defect, an incorrect change, or a test that never ran. If a defect got through, did it get through because of human error? Their historical range is that between two-thirds and four-fifths of major failures involve some human element.
So "human error" here does not mean somebody was careless. It means a person was somewhere in the chain, which is true of almost everything an operation does. The useful detail is not that people are involved. It is how.
Asked how human error contributed to outages over the past three years, operators could pick up to three answers (n=199). Two dominate.
Uptime Institute Data Center Resiliency Survey 2026, n=199. Respondents could select up to three answers, so these overlap and do not sum to 100. Bars are scaled to the largest response, not to a share of a whole.
Look at the top two.
Fifty-nine per cent: the procedure existed and was not followed.
Thirty-six per cent: the procedure was followed, and it was wrong.
Most organisations respond to both with the same two actions. Retrain the person. Reissue the document.
For the first, that is a reasonable start. For the second it is worse than useless. It reinforces a procedure that has already been shown not to work, and it tells the person who hit the problem that the fault was theirs.
These are different failures and they need different responses. The first asks why following the procedure was harder than not following it. The second asks how nobody noticed the procedure had stopped matching the job.
A procedure that one person skips once is an incident. A procedure that gets quietly worked around by experienced staff, repeatedly, without anyone raising it, is not really an adherence problem. It is a design problem wearing an adherence problem's clothes.
The people closest to the work usually know exactly which procedures those are. They rarely put it in writing, because the written process is the one that will be quoted back at them if something goes wrong.
Uptime's own conclusion points the same way: managing human error is less about eliminating mistakes and more about designing systems that anticipate them.
The pattern is not specific to data centres. Here is what the split looks like elsewhere.
Every gate is passed and the change still fails weeks later. Often the rollback step was documented but never rehearsed under real load, or the test environment did not carry production volumes. The procedure was followed exactly. It described a system that no longer exists.
A control is signed off, but the check behind it spans four systems and twenty minutes at month end. It gets signed on trust. That is not a training problem. The control was designed for a volume that no longer exists.
A picking or loading step is skipped at peak because following it fully makes the shift target unreachable. When the procedure and the target contradict each other, the target wins every time. The procedure is not the thing being managed.
Agents use a workaround that is faster than the documented flow and produces a better outcome for the customer. Nobody reports it, because raising it looks like admitting a breach. The knowledge stays undocumented until the person who holds it leaves.
A maintenance interval was set years ago against different equipment or a different duty cycle. It is followed precisely. It is still wrong.
In four of those five, retraining the person changes nothing.
We want to know what actually happens in operations teams, not what the framework says should happen. Results go into a later issue. One click, no sign-in.
One vote per browser. No sign-in, nothing stored about you beyond the count.
Ask them in this order. The first splits the problem in two, and the two halves need different fixes.
One addition worth making permanent. Most operations have a route for reporting incidents. Very few have a route for reporting a procedure that does not work, before it causes one. Give people that route, and make clear that using it is not an admission of anything.
Worth being straight about, because these numbers get repeated carelessly elsewhere.
This is survey data from data centre operators. The Data Center Resiliency Survey 2026 was conducted in February and March 2026 with more than 450 operator respondents. The Global Data Center Survey 2025 was conducted in the second quarter of 2025 with more than 800 operators. Individual questions have much smaller bases, stated above.
It is self-reported. The 87 per cent preventability figure is what operators believe about their own failures, not an independent finding, and people assessing their own incidents may lean either way.
The cause questions allowed up to three answers, so the percentages overlap. Fifty-nine and 36 are not two halves of a whole.
And it is one sector. Data centres are not banking operations or warehouses. What transfers is not the percentage. It is the distinction between a procedure that was skipped and a procedure that was wrong, and the fact that most organisations only have a response for the first one.
Neither is about outages. Both are about the gap between a documented process and the one people actually run. Video content can be audited free on both; assessments and the certificate sit behind the paid tier. Content and pricing change, so check before you enrol.
Eighty-seven per cent preventable sounds like an indictment. It is closer to the opposite.
If most failures came from things outside your control, there would be very little to do beyond buying more redundancy and hoping. They do not. They come from management, process and configuration, which are three things an operations team can change without a budget round.
The organisations that get better at this are not the ones with the strictest procedures. They are the ones that treat a skipped procedure as information rather than as a disciplinary matter, and that check their written processes against the work often enough to catch the ones that quietly stopped being true.
Keep Sharing... Keep Learning...
Sandeep
The OperationsCareers Team
OperationsCareers is building a focused publication for professionals working across operations, transformation, process improvement, PMO, reporting and workflow optimisation.
If you are hiring operational talent, promoting a relevant tool or course, or looking to reach a practical professional audience, we are open to selective sponsorship and featured opportunities.
The views expressed in this newsletter are based on personal professional experience and observation. They do not constitute professional, legal, financial or career advice.