Open App →
Back to Blog
September 4, 2026 By Steve

Redundant at the Wrong Layer

...redundancy is only ever redundant against the specific failure it was designed for...

Steve Williams Babnews.org

For a stretch yesterday morning, most of the country's AI assistants stopped answering at the same time. ChatGPT, Claude, and Grok all went dark inside the same window. Google's Gemini kept running. [verify: durations and the exact list of affected services are still being confirmed. OpenAI's disruption looks shorter, on the order of half an hour, while Anthropic's ran longer, closer to three hours.]

The question everyone asked was the right one. These are among the best-funded, most technically capable companies on the planet. They employ people who think about failure for a living. So why didn't redundancy save them?

It did. That's the part worth sitting with. The redundancy was there. It just lived one layer above the layer that failed.

Give the engineering its due

Start by taking these companies seriously, because the lazy version of this story is that they were careless, and they weren't. Every major AI provider runs across multiple availability zones. They run failover. They load balance. They replicate. If a single server rack or a single building goes down, traffic reroutes and most users never notice. That machinery is real, and on an ordinary bad day it works.

But redundancy is only ever redundant against the specific failure it was designed for. Duplicate your servers and you're covered against a server dying. You are not covered against the thing underneath the servers dying. And yesterday the thing that failed was underneath.

The layer that bound

Reports disagree on the exact culprit. One account points to a compute center in Memphis. Another points to a cloud region on the east coast and a shared control plane inside it. [verify] The disagreement doesn't matter for the mechanism, and that's the tell. Whatever the specific component was, it was a piece of physical infrastructure that more than one supposedly independent AI company was quietly standing on.

When the fault is a shared region, a shared network backbone, a shared control plane, or a shared compute partner, your three redundant zones stop being three. They become one, because they all fall through the same trapdoor at the same instant. Redundancy at the application layer does nothing about a fault two layers down. You built three doors out of the room, and all three opened onto the same collapsing floor.

Why real redundancy is close to impossible right now

Here's where it stops being a software problem and becomes the problem I keep coming back to.

To be genuinely redundant at the frontier, you would need a second stack that shares nothing with the first. A different region, a different power feed, a different network path, a different pool of accelerators. Not a different logo on the invoice, a physically independent duplicate.

That requires slack, and the buildout has none. The scarcest inputs to frontier AI are the advanced-packaging capacity and high-bandwidth memory that gate the chips, the data-center space that's gated by power, and the grid interconnects that are gated by transformers and permitting. All of it is rationed. A fully independent backup competes for the exact same rationed supply that growth needs. And in a land-grab phase, when every competitor is racing to capture the market, growth wins that fight every time. Resilience is a cost with no revenue attached and a payoff you only see on the rare day it saves you. So it gets underbuilt, right up until an outage puts a price on it.

This is why the intelligence doesn't rescue them. AI is very good at compressing the bottlenecks you can write down as information. It is close to useless against the bottlenecks made of physics. No model, however capable, conjures a second independent power feed, a second packaging line, or a second region that isn't already spoken for. The intelligence is real. It's just pointed at a layer that wasn't the one that broke.

Diversified at the logo level, concentrated at the physical level

The industry told itself a comforting story, and the story was true at the layer it described and false at the layer that mattered.

At the model layer, the market looks diverse. Different companies, different architectures, different clouds. But those clouds rest on a short list of the same physical dependencies. The same few regions absorb a large share of the traffic. The same grids feed them. In some cases the same merchant compute supplies more than one of them, which is what the Memphis account implies: rivals leaning on a shared source of capacity without either set of users knowing. [verify]

So the diversification was real on the map and thin on the ground. The map showed many independent providers. The territory was one substation. This is the same shape as the concentration I've written about on the money side of the buildout, where a single keystone sits at the center of everyone's capital and control stops requiring ownership. Here it's the same geometry in copper and concrete: control of the layer everyone shares, whether or not anyone set out to build it that way.

The backup was correlated with the failure

There's a final twist that turns a bad outage into a synchronized one. Redundancy assumes failures are independent. Yesterday they were the opposite of independent.

The moment ChatGPT went down, its users didn't wait. They poured into Claude and Grok. So the systems that were supposed to serve as the market's informal backup took a demand spike at the precise moment the shared infrastructure beneath them was already degraded. The fallback got hit twice, once by the underlying fault and once by everyone else's refugees. You can't fail over to the system that's about to eat your traffic while you're both standing on the same sick floor.

Independent redundancy would have absorbed that. Correlated redundancy amplified it.

What the outage actually revealed

Markets had been pricing these services as diversified, resilient, utility-grade infrastructure. That's the growth read, and it lives on one axis. The outage exposed the other axis, the control side, the one that stays invisible until it fails. The tell was there the whole time, in the shared regions and the shared providers, on nobody's dashboard because nothing had forced it onto one.

So the honest conclusion isn't that these companies forgot to build redundancy. It's that you cannot be redundant against a constraint you all depend on. You can only avoid depending on it, and avoiding it costs capacity nobody is willing to spend while the market is still being captured.

You don't remove a bottleneck. You relocate it. Yesterday we learned that the whole industry, racing in what looked like different directions, had relocated its single point of failure onto the same rung.

The question worth carrying out of it isn't about AI at all. It's this: where else are we calling something diversified when only the label is?