Availability Blind Spots
The availability you think you have, and the availability you actually have. Every article here is a risk that did not look like a risk: an outage you scheduled yourself, a login you do not own, a single architect, and a priority list where everything is critical.
-
Planned Outages are Still Outages
Customers don't care whether an outage was planned. Why routine maintenance windows quietly cap your real availability far below the 99.9%+ your team thinks it's hitting, and why the fix is zero-downtime deployment, not a bigger excuse.
Read more → -
The Hidden Risks of "Sign In with Google"
"Sign In with Google" is convenient, but it turns your Google account into a single point of failure — locked out of Google means locked out of everything linked to it. Why password managers and passkeys give the same one-click convenience without the coupling.
Read more → -
The Single-Architect Availability Problem
Reliability engineering asks what happens when a component becomes unavailable. Most organizations never ask that about the person who holds all the architectural context — and the six-week test to find out if you have that single point of failure.
Read more → -
When Everything Is Critical, Nothing Is
Reliability isn't only a technical property, it's an ownership property. Two decisions sit underneath every reliable system, what actually matters and who is accountable for keeping it up, and most organizations have made neither.
Read more →