Yesterday's Amazon Web Services (AWS) region outages is a reminder that when Amazon sneezes, the internet catches a cold. Time for reflection over immediate rebuild discussions.
The region outage in Northern Virginia sparked quite some noise. Unsurprisingly the internet is flooded with opinions on how dependent Europe has become on U.S. hyperscalers. Including the risk of application deployment in just a single region and the need for ‘European control’ over our digital infrastructure.
To be fair, people have a point. Europe’s internet is heavily dependent on Amazon, Microsoft, and Google. A full AWS region outage, with multiple availability zones should not happen, period! Yet, we shouldn’t overreact either.
From my perspective, the EU dependency on US service providers isn’t new at all; it’s more of a shift. Ten - Fifteen years ago, when Akamai or another large CDN had a global outage, the same pattern appeared. Popular internet services became unreachable worldwide. And let’s not forget the transatlantic cables connecting Europe and the U.S. From the first telegraph cable constructed by Cyrus Field in 1855 to today’s rapidly expanding fiber optic network we see that the majority of these construction initiatives are led & owned by American companies.
That doesn’t mean European initiatives aren’t valuable, they absolutely are, even if we’re a decade late to the game. But the reflex we’ve seen in the past 24 hours, to rebuild everything ourselves and deploy all applications across multiple regions or perhaps in our own data centers, is a bit simplistic.
Regional outages (or in case of on-premises, full data center failures) are rare. Building multi-region or multi-datacenter architectures requires significant investment and puts heavy pressure on DevOps teams. Combined that in real life, most outages come from the way we design workloads, rather than platform wide failures.
Multi-region architectures make sense for mission-critical workloads, in practice, just a small subset of your ecosystem. A more data driven approach means weighing the cost of downtime against the likelihood and impact of those incidents compared to the incremental costs of reducing recovery time.
Sometimes the key lesson from an outage isn’t how to stop AWS from ever failing again, but to better understand and own our dependencies.
Reflection instead of reflex!