Issue #10 · The Ops & Transformation Briefing · September 2026

The equipment held. The procedure didn't.

Nine in ten major IT failures involve human error. Almost none involve carelessness — and the fix is more often the procedure than the person.

Issue #10 8 min read Reader poll Free preview
This issue in one line

When something fails, most organisations fix the person. The data says the procedure is just as likely to be the problem.

A change goes through every stage. Requirement, design, testing, approval, release. Every step is signed. It goes to production and works. A month later it fails.

That sequence will be familiar to anyone who has run a change process. Nobody cut a corner. Nobody was careless. Every gate was passed by someone competent doing what the process asked. The outcome was still a failure, arriving late enough that the connection to the change was not obvious at first.

Then comes the review, and the finding is almost always the same, and almost always vague: update the procedures. Which procedures, updated how, and why they were wrong in the first place, rarely survives into the action log.

A worn procedure manual mounted on a wall, with a newer handwritten note taped up beside it

Most organisations, faced with a failure, do two things. Retrain the person. Reissue the document.

The Uptime Institute has been asking operators how human error actually contributes to their failures. The answers split into two groups, and they are not the same problem. One of them responds to retraining. The other gets worse when you try it.

If the people involved were competent and following the process, what exactly is failing?

Read the full Issue 10 briefing

Free, every other Friday. Issue 10 has the figure for how many failures operators say were preventable, the two causes that get treated as one problem, why a procedure that keeps getting skipped is telling you something, worked examples across change and release, finance operations, logistics, customer operations and maintenance, three questions to ask after the next failure, and an honest account of what this data does and does not show.

After confirming your email, you'll be taken to the subscriber download page. Unsubscribe any time.

One question, whether you subscribe or not

We want to know what actually happens in operations teams, not what the framework says should happen. Results go into a later issue. One click, no sign-in.

When something goes wrong in your operation, what usually happens first?

One vote per browser. No sign-in, nothing stored about you beyond the count.

Two courses on process that survives contact with the work

Neither is about outages. Both are about the gap between a documented process and the one people actually run. Video content can be audited free on both; assessments and the certificate sit behind the paid tier. Content and pricing change, so check before you enrol.

Foundation · Process design
Introduction to Operations Management — University of Pennsylvania (Wharton)
Coursera · Around four weeks · Self-paced
Process analysis, and the reasons designed capacity and real capacity diverge. Useful background if your corrective actions keep addressing behaviour rather than the process that shapes it.
View on Coursera →
Practical · Improvement method
Six Sigma and Lean: Quantitative Tools for Quality and Productivity — University System of Georgia
Coursera · Around four weeks · Self-paced
Root cause analysis done properly, which is the discipline missing when a review concludes "update the procedures" without saying which one or why.
View on Coursera →