Designing systems that survive component failures, recovering data and service when they don't, and learning from every incident.
- Failure Domains and Single Points of Failure
- Finding the components, zones and credentials whose loss takes everything down.
- Redundancy and Failover
- Replicas, multiple zones and automatic failover, and testing that failover actually works.
- Graceful Degradation
- Circuit breakers, fallbacks and load shedding, so a failing feature doesn't take down the whole product.
- Backups and Restore Drills
- Backups on a schedule, stored separately from the original, and restored regularly to prove they work.
- Recovery Objectives
- Recovery point and recovery time objectives that state how much data loss and downtime the business can accept.
- Incident Response and Postmortems
- Mitigating first, communicating clearly, and turning each incident into new tests, alerts and gates.