>...one of the standard steps is to shift traffic off of one of the redundant routers in the primary EBS network to allow the upgrade to happen. The traffic shift was executed incorrectly...
This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.
This supports the theory that between 50%-80% of outages are caused by human error
Not quite - in this case, a single human error then triggered a series of latent and undiscovered bugs in the system itself. It's a confluence of small events that makes for a large-scale problem like this.
Unfortunately the humans in reference here are supposed to be the best of the lot. This brings in the eternal question during times of crisis ,how are the best very different than the mediocre or the average?
If not then with little hard work and smart work here and there any body can beat these 'best' during non crisis times. And during the crisis time all are same any way.
Probably that's why there are a lot of successful companies even with average talented people.
Any high availability system inherently aims for that, but "never" is a strong word. Sooner or later, you need to upgrade or replace physical hardware.
"The trigger for this event was a network configuration change. We will audit our change process and increase the automation to prevent this mistake from happening in the future."
Why they didn't do what? "Increase the automation"? I suspect that they have been doing that from the beginning. It's an ongoing process.
I hate to quote Rumsfeld, but there are known unknowns, and unknown unknowns. Of course you want to eliminate the latter-- but there's (necessarily) no way you can ever know that you've done so.
This supports the theory that between 50%-80% of outages are caused by human error, regardless of the resilience of the underlying infrastructure.