One morning last October — the 20th — a piece of Amazon’s cloud in Northern Virginia lost track of its own address. An internal system whose whole job is to keep a database reachable pointed everyone at nothing — an empty entry where a location should be. For about fifteen hours, that one fault rolled outward through more than a thousand services. Snapchat, Signal, Slack, Robinhood, Coinbase, Fortnite, Alexa — even the UK tax office. Downdetector logged over six million complaints in a day.
Nobody attacked anything. Amazon’s own automation made a routine change, the change was wrong, and the wrongness travelled at the speed of software.
The comforting story is that this was a freak event. It wasn’t. It was a landlord problem — and almost every business online has the same landlord.
Three landlords
Strip away the branding and the entire web rents from three companies. Amazon holds roughly 30% of the cloud market, Microsoft about 20%, Google around 13%. Together, close to two-thirds of everything.
Your favourite app, your bank’s login page, the shop you bought shoes from last week — most of them don’t own the computers they run on. They rent, and they rent from the same short list. When people say “the cloud”, they mean one of three buildings.
Concentration is the whole pitch, and the pitch is honest. One provider, run by specialists, is cheaper and steadier than a server humming in your own office cupboard. That’s true — right up until the day it isn’t, and then it isn’t true for everyone at once.
Which means two sentences a business quietly tells itself are really one. “We run on the biggest, safest provider” and “we go down at the exact moment everyone else does” describe the same decision.
The address book. Amazon has one region in Northern Virginia that is older and bigger than the rest, and a surprising amount of its global plumbing is anchored there — the systems that log you in, that turn a web address into a location. Four of Amazon’s five worst outages in recent years started in that one place. Rent a flat across town and you’re still on the same building’s water main.
One bad day
Here’s the part that should change how you think about it. Line up the big outages of the past two years and they rhyme.
Google, June last year. A software update carried a blank field where a number should be, and it crashed a core service worldwide, over and over. Seventy-odd Google products went down, and Spotify and Discord for company. It reached further, too: Cloudflare — the service many companies pay to stay online when their main cloud stumbles — had quietly parked some of its own plumbing on Google, and went down with it. The safety net shared a landlord with the thing it was protecting.
Amazon, last October. The empty address above. Fifteen hours, a thousand services.
Microsoft, nine days after Amazon. A bad configuration change slipped past a safety check and spread across its global front door. Office, Teams, Xbox — plus Costco, Starbucks, and Alaska Airlines. Eight hours.
Cloudflare, last November. A routine database tweak quietly doubled the size of an internal file until it hit a hard limit, and the software fell over. X, ChatGPT, Spotify. Five and a half hours.
None of these were attacks. Every one was self-inflicted — a small, routine change, made by competent people, that turned out to be wrong and then copied itself everywhere before anyone could pull it back. The same machinery that makes the cloud cheap and uniform is the machinery that makes it fail uniformly. When every server runs the identical config, one bad config takes down every server.
You don't rent a safer internet by paying a bigger landlord. You rent the same bad Tuesday everyone else is having.
Price it like weather
So what does a business actually do with this?
Start by dropping the fantasy that there’s a safer provider to switch to. There isn’t. All three have taken themselves offline with a typo in the last eighteen months. Reliability shopping is choosing which coin to flip.
Read the fine print on “99.99%”. That famous number sounds like a promise. It permits about fifty minutes of downtime a year, and when a provider sails past it, the compensation is a credit against your next bill — not a cheque for what the outage cost you. You have to apply for it. It’s a coupon, not insurance.
Cost the outage before it happens. For most mid-sized companies, an hour offline runs past $300,000 once you count lost sales, idle staff, and the support queue. The honest planning question was never “how do we never go down” — you will — it’s “what does our business do during the hours we’re down, and have we ever once rehearsed it?”
Keep a version that survives. The cheapest resilience isn’t a second cloud; it’s a simpler front door. A plain, static copy of your most important pages — a menu, a phone number, an address, a “we know, we’re on it” — can sit somewhere that doesn’t share your main provider’s bad day. When the clever machinery falls over, the boring page stays up.
The cloud didn’t sell you fragility. It sold you a genuinely good deal, and the deal is real. But the bill for it arrives all at once, on a day you don’t get to choose, alongside everyone else who signed the same lease.
You can’t pick a landlord who never has a bad day. You can decide, before it comes, what your business looks like on theirs.