Availability & Resilience
SLAs and the gap between contractual and achievable, failure modes, graceful degradation, blast radius, criticality tiers, service ownership, and dependency mapping.
-
It Passed the Test. That Doesn't Mean It Works.
Every testing, monitoring, and reliability practice you own rests on one assumption: same input, same output. An LLM in the call path removes it. What replaces it is evaluation, thresholds, and verification as an explicit architectural layer with a cost and an owner.
Read more → -
When Everything Is Critical, Nothing Is
Reliability isn't only a technical property, it's an ownership property. Two decisions sit underneath every reliable system, what actually matters and who is accountable for keeping it up, and most organizations have made neither.
Read more → -
The Single-Architect Availability Problem
Reliability engineering asks what happens when a component becomes unavailable. Most organizations never ask that about the person who holds all the architectural context — and the six-week test to find out if you have that single point of failure.
Read more → -
The Hidden Risks of "Sign In with Google"
"Sign In with Google" is convenient, but it turns your Google account into a single point of failure — locked out of Google means locked out of everything linked to it. Why password managers and passkeys give the same one-click convenience without the coupling.
Read more → -
Planned Outages are Still Outages
Customers don't care whether an outage was planned. Why routine maintenance windows quietly cap your real availability far below the 99.9%+ your team thinks it's hitting, and why the fix is zero-downtime deployment, not a bigger excuse.
Read more →