
On November 18, 2014, Microsoft Azure went dark. Thousands of the cloud computing service’s customers experienced downtime on their sites for over 9 hours. When they flooded Microsoft customer support to ask what the hell had gone wrong, customers learned that it wasn’t some glitch, natural disaster, or devious hacking scheme.
It was pure human error.

Microsoft deployed an update without running through the standard operating guidelines specifically laid out for this scenario. Instead of rigorously checking that the update was good to go, engineers shipped it on the assumption it was bug-free. This wasn’t just this one-time incident—engineers at Azure regularly violated standard operating procedure because 99.9% of the time, it was a total waste of their resources.
And it’s not just Azure that does this—habitual violation of process leads to a ton of mistakes in all kinds of IT arenas.
When I was 16, I landed my first “real” job which so happened to be in IT. I had just passed my CCNA with the help of my dad (apparently the youngest person in Australia at the time) and was hungry to get my hands on some real technology.
But once I started, I found there was rigid process everywhere, and configuring a live Windows 2000 server was a totally different experience to tinkering with my fathers machines at home, with serious consequences for not following the process.
Continue Reading









