Why 100% Availability Is the Wrong Target?
100% availability. Of course it’s something everyone wants, would love to have, dreams about. But the reality is that it’s only beautiful on paper. Just like everything that’s perfect, it looks great in theory and falls apart the moment it meets production.
Here’s the uncomfortable truth. If you actually chase 100% availability as a target, not close to 100% but literally 100%, you invite a whole set of problems. I know because I lived them.
The Fear Culture I Used to Live In
Back when I worked as a sysadmin (~2015), I carried the belief that I could not fail. I could not make a mistake. There was no room for error.
That mindset had consequences. Every meaningful action had to happen in one specific window, late at night or on a weekend, and me and a thousand other people were terrified every single time. The fear didn’t make us safer. It made everything slow.
Every new feature or patch we shipped required a small window of unavailability. Sure, there are plenty of techniques to minimize that. But culturally, the “everything must be up 100% of the time” idea simply doesn’t work. It breeds a team that’s too scared to deliver, and a team afraid to ship is, without a doubt, a risk in itself. So the team never bring new ideas.
This is exactly what SRE talks about. You define a good SLO, you set an error budget, and you shift the culture. People stop obsessing over failing and start thinking about how to minimize the impact of failure.
Focus on Impact, Not on Never Failing
Here’s the mental shift that changed everything for me. The goal isn’t to reduce the number of failures. It’s to reduce their impact.
Think about it. Keeping something down for one hour, yes, I agree, that’s a lot. That’s a real problem. But that statement quietly assumes you’re affecting 100% of your customers.
What if that one hour only affects 1% of your customers? And what if not all of them are even trying to access the system at that moment? Suddenly the same “one hour of downtime” is a completely different conversation.
That’s the reframe I want to share today. Downtime is not automatically 100% impact. When you stop measuring reliability as “were we perfectly up?” and start measuring it as “how many users actually felt this, and for how long?”, your whole engineering posture changes.
This is why SLOs and error budgets exist. They give you permission to take risks intelligently:
- 99.9% availability ≈ 43 minutes of budget per month
- Spend that budget on canary releases, feature rollouts, and experiments
- When you’re burning it too fast, that’s your signal to slow down, not fear, not guesswork
The error budget turns “don’t fail” into “here’s exactly how much risk we can afford this month.” That’s empowerment, not permission to be careless.
Also, at least talking about error budget empower the team to think on “HOW CAN I SAVE THAT BUDGET OR MAKE IT BETTER ?”
You Don’t Control 100% of the Chain Anyway
There’s another brutal, often ignored factor. Even if you achieved a flawless culture and swore you’d never fail, you still depend on things that are themselves below 100%.
Take a mobile app. Your user is on 4G or 5G. How many things can go wrong on that path that have absolutely nothing to do with your application? The network drops, the tower is congested, their signal fades in an elevator. To the user it feels like your app failed, but you never even received the request.
So even if you spent a fortune building a “never fails, never will” culture, it doesn’t matter. The other teams, the third-party APIs, the DNS provider, the payment gateway, the cloud region, none of them are giving you a 100% SLA either.
Reliability is a chain. Your availability can never exceed the combined availability of everything you depend on. Chasing 100% for your slice while your dependencies deliver 99.9% is chasing a number that physically cannot exist.
What This Looks Like in Practice
So what do you do instead? A few practical moves that can help you:
- Define SLOs per service based on real user impact, not a blanket “100% everywhere.” or best-effort 100%
- Measure user-centric SLIs like “% of successful requests” or “% of users served within latency target,” not just “was the box up?”
- Set a realistic error budget and let teams spend it on shipping features safely.
- Reduce blast radius with canary releases, feature flags, and gradual rollouts so a bad change hits 1% of users, not 100%.
- Alert on burn rate, not vanity metrics. “We’ll exhaust this month’s budget in 2 hours” is actionable. “CPU is high” is noise.
The Takeaway
100% availability is beautiful in theory and toxic as a target. It creates a culture of fear, slows delivery pace, ignores the dependencies you don’t control, and measures the wrong thing entirely.
The mature move, the SRE move, is to accept that failure is inevitable, define what “good enough” actually means for your users, and spend your error budget building better systems instead of hiding from imperfect ones.
Stop optimizing for never failing. Start optimizing for minimal impact when you do.
If you want to keep getting quick tips like this, the everyday things every SRE should be thinking about, stay subscribed to the newsletter. There’s plenty more coming.
Cheers,
Douglas Mugnos
MUGNOS-IT 🚀