Issue #10 · The Ops & Transformation Briefing · September 2026

The equipment held. The procedure didn't.

Nine in ten major IT failures involve human error. Almost none involve carelessness — and the fix is more often the procedure than the person.

Issue #10 8 min read Reader poll Full briefing
This issue in one line

When something fails, most organisations fix the person. The data says the procedure is just as likely to be the problem.

A change goes through every stage. Requirement, design, testing, approval, release. Every step is signed. It goes to production and works. A month later it fails. The review concludes that the procedures need updating.

That sequence will be familiar to anyone who has run a change process. Nobody cut a corner. Nobody was careless. Every gate was passed by someone competent doing what the process asked. The outcome was still a failure, arriving late enough that the connection to the change was not obvious at first.

The finding is almost always the same, and almost always vague: update the procedures. Which procedures, updated how, and why they were wrong in the first place, tends not to survive into the action log.

There is now reasonable data on how often this happens and what sits behind it. Start with the encouraging part, because it is real.

Failures are becoming less common. In the Uptime Institute's Global Data Center Survey 2025, half of operators said they had experienced an impactful outage in the previous three years. That is the lowest figure recorded since 2020, when the same question returned 74 per cent.

Operators reporting an impactful outage in the previous three years
202074%
202169%
202260%
202355%
202453%
202550%

Uptime Institute Global Data Center Survey 2025. Sample sizes by year: 494, 730, 730, 781, 768, 754. Five consecutive years of improvement.

A worn procedure manual mounted on a wall, with a newer handwritten note taped up beside it

Investment in equipment, monitoring and process has worked. That is the story of the last five years, and it deserves saying before the rest.

Now the part worth sitting with. Among operators who had an impactful outage in the past three years, 87 per cent said it could have been avoided with better management, processes or configuration. That is up seven percentage points on the previous year. It is a small base — 98 respondents — and it is their own judgement rather than an independent assessment.

Nearly nine in ten failures were inside somebody's control.

Read that the pessimistic way and it is damning. Read it the useful way and it is the most encouraging number in the report. Not weather, not fate, not a supplier. Management, process, configuration.

Which raises the obvious question. If these failures were preventable, and the people involved were competent and following the process, what exactly is failing?

Nine in ten failures have a person in them. Almost none have a villain.

In the Uptime Institute Data Center Resiliency Survey 2026, 92 per cent of respondents said human error contributed at least something to their most recent significant outage (n=220). Thirty-one per cent called it a major contributor, 30 per cent moderate, 31 per cent minor. Only 8 per cent said it contributed nothing.

That number gets repeated a lot, usually badly. So note what Uptime itself says about it. They do not count human error as a root cause. They treat it as a contributing factor, because attribution is genuinely hard. A failed update might be a software defect, an incorrect change, or a test that never ran. If a defect got through, did it get through because of human error? Their historical range is that between two-thirds and four-fifths of major failures involve some human element.

So "human error" here does not mean somebody was careless. It means a person was somewhere in the chain, which is true of almost everything an operation does. The useful detail is not that people are involved. It is how.

Two different problems that get treated as one

Asked how human error contributed to outages over the past three years, operators could pick up to three answers (n=199). Two dominate.

How human error contributed to outages
Staff failed to follow procedures59%
The procedures themselves were wrong36%
Installation issues25%
In-service issues25%
Insufficient staff23%
Preventative maintenance frequency16%
Design issues or omissions16%

Uptime Institute Data Center Resiliency Survey 2026, n=199. Respondents could select up to three answers, so these overlap and do not sum to 100. Bars are scaled to the largest response, not to a share of a whole.

Look at the top two.

Fifty-nine per cent: the procedure existed and was not followed.

Thirty-six per cent: the procedure was followed, and it was wrong.

Most organisations respond to both with the same two actions. Retrain the person. Reissue the document.

For the first, that is a reasonable start. For the second it is worse than useless. It reinforces a procedure that has already been shown not to work, and it tells the person who hit the problem that the fault was theirs.

These are different failures and they need different responses. The first asks why following the procedure was harder than not following it. The second asks how nobody noticed the procedure had stopped matching the job.

A procedure that gets skipped often is telling you something

A procedure that one person skips once is an incident. A procedure that gets quietly worked around by experienced staff, repeatedly, without anyone raising it, is not really an adherence problem. It is a design problem wearing an adherence problem's clothes.

The people closest to the work usually know exactly which procedures those are. They rarely put it in writing, because the written process is the one that will be quoted back at them if something goes wrong.

Uptime's own conclusion points the same way: managing human error is less about eliminating mistakes and more about designing systems that anticipate them.

That is not a soft observation. It is the difference between a corrective action that holds and one that quietly fails again in eight months.

The same two questions, five different operations

The pattern is not specific to data centres. Here is what the split looks like elsewhere.

1

Change and release

Every gate is passed and the change still fails weeks later. Often the rollback step was documented but never rehearsed under real load, or the test environment did not carry production volumes. The procedure was followed exactly. It described a system that no longer exists.

2

Financial services operations

A control is signed off, but the check behind it spans four systems and twenty minutes at month end. It gets signed on trust. That is not a training problem. The control was designed for a volume that no longer exists.

3

Logistics and warehouse

A picking or loading step is skipped at peak because following it fully makes the shift target unreachable. When the procedure and the target contradict each other, the target wins every time. The procedure is not the thing being managed.

4

Customer operations

Agents use a workaround that is faster than the documented flow and produces a better outcome for the customer. Nobody reports it, because raising it looks like admitting a breach. The knowledge stays undocumented until the person who holds it leaves.

5

Maintenance

A maintenance interval was set years ago against different equipment or a different duty cycle. It is followed precisely. It is still wrong.

In four of those five, retraining the person changes nothing.

One question before you carry on

We want to know what actually happens in operations teams, not what the framework says should happen. Results go into a later issue. One click, no sign-in.

When something goes wrong in your operation, what usually happens first?

One vote per browser. No sign-in, nothing stored about you beyond the count.

Three questions to ask after the next failure

Ask them in this order. The first splits the problem in two, and the two halves need different fixes.

After the next failure
  1. Was the procedure followed? Answer this before anything else, and answer it honestly rather than diplomatically.
  2. If it was not followed — what made not following it easier? Time pressure, tooling, a target that conflicts with it, a step that cannot be completed as written. Look for the incentive before you look for the individual.
  3. If it was followed — when was it last checked against how the work is actually done? Not reviewed for compliance. Checked against reality, with the people doing it.

One addition worth making permanent. Most operations have a route for reporting incidents. Very few have a route for reporting a procedure that does not work, before it causes one. Give people that route, and make clear that using it is not an admission of anything.

What this data is, and what it is not

Worth being straight about, because these numbers get repeated carelessly elsewhere.

This is survey data from data centre operators. The Data Center Resiliency Survey 2026 was conducted in February and March 2026 with more than 450 operator respondents. The Global Data Center Survey 2025 was conducted in the second quarter of 2025 with more than 800 operators. Individual questions have much smaller bases, stated above.

It is self-reported. The 87 per cent preventability figure is what operators believe about their own failures, not an independent finding, and people assessing their own incidents may lean either way.

The cause questions allowed up to three answers, so the percentages overlap. Fifty-nine and 36 are not two halves of a whole.

And it is one sector. Data centres are not banking operations or warehouses. What transfers is not the percentage. It is the distinction between a procedure that was skipped and a procedure that was wrong, and the fact that most organisations only have a response for the first one.

Two courses on process that survives contact with the work

Neither is about outages. Both are about the gap between a documented process and the one people actually run. Video content can be audited free on both; assessments and the certificate sit behind the paid tier. Content and pricing change, so check before you enrol.

Foundation · Process design
Introduction to Operations Management — University of Pennsylvania (Wharton)
Coursera · Around four weeks · Self-paced
Process analysis, and the reasons designed capacity and real capacity diverge. Useful background if your corrective actions keep addressing behaviour rather than the process that shapes it.
View on Coursera →
Practical · Improvement method
Six Sigma and Lean: Quantitative Tools for Quality and Productivity — University System of Georgia
Coursera · Around four weeks · Self-paced
Root cause analysis done properly, which is the discipline missing when a review concludes "update the procedures" without saying which one or why.
View on Coursera →

The good news is the whole point

Eighty-seven per cent preventable sounds like an indictment. It is closer to the opposite.

If most failures came from things outside your control, there would be very little to do beyond buying more redundancy and hoping. They do not. They come from management, process and configuration, which are three things an operations team can change without a budget round.

The organisations that get better at this are not the ones with the strictest procedures. They are the ones that treat a skipped procedure as information rather than as a disciplinary matter, and that check their written processes against the work often enough to catch the ones that quietly stopped being true.

Keep Sharing... Keep Learning...

Sandeep
The OperationsCareers Team

Source note. The figures in this issue come from Uptime Institute, Annual outage analysis 2026 (Keynote Report 201, May 2026), by Douglas Donnellan, Andy Lawrence and Rose Weinschenk. It is the eighth edition of an annual series. The report draws on three sources: the Uptime Institute Data Center Resiliency Survey 2026, conducted in February and March 2026 with more than 450 operator and 500 vendor respondents; the Uptime Institute Global Data Center Survey 2025, conducted in the second quarter of 2025 with more than 800 operator respondents; and a database of publicly reported outages, of which 122 were recorded in 2025. Sample sizes for individual questions are given alongside each figure above and are considerably smaller than the survey totals. Uptime states that it treats human error as a contributing factor rather than a root cause, and that its public outage database is directionally useful rather than representative. All figures are self-reported by operators.

The views expressed in this newsletter are based on personal professional experience and observation. They do not constitute professional, legal, financial or career advice.