Building Software

Engineering Fundamentals for the Agent Era

Contents Section 7, Operations

Reliability, Redundancy and Recovery

Mistakes to catch in review

  1. Backups that were never restored and turn out to be empty, partial or corrupt when they are needed.

  2. Two 'redundant' instances that share one database, one availability zone or one credential.

  3. A health check that reports healthy while the application can't reach its database.

  4. An incident closed the moment the symptom stops, before the cause is known, so it comes back the next week at the same hour.

Designing systems that survive component failures, recovering data and service when they don't, and learning from every incident.

Topics

Failure Domains and Single Points of Failure
Finding the components, zones and credentials whose loss takes everything down.
Redundancy and Failover
Replicas, multiple zones and automatic failover, and testing that failover actually works.
Graceful Degradation
Circuit breakers, fallbacks and load shedding, so a failing feature doesn't take down the whole product.
Backups and Restore Drills
Backups on a schedule, stored separately from the original, and restored regularly to prove they work.
Recovery Objectives
Recovery point and recovery time objectives that state how much data loss and downtime the business can accept.
Incident Response and Postmortems
Mitigating first, communicating clearly, and turning each incident into new tests, alerts and gates.

You understand it when you can

  • Map the single points of failure in a system diagram and propose a fix for the worst one.
  • Restore a service's data from backup and time how long it takes.
  • Write a blameless postmortem for an outage, with a timeline and at least one systemic fix.

Drill

An agent set up nightly database backups to a storage bucket in the same account and region as the database, plus a health check endpoint that returns 200 without touching the database. Find the failure scenarios in which both mechanisms report success while the service or its data is gone.

Start here

Watch

Velocity 2012: Richard Cook, "How Complex Systems Fail"

Richard Cook, 2012. 28-minute talk.

Cook explains why complex systems always run partly broken and why 'human error' is a starting point for an investigation rather than a root cause, which is the basis of blameless postmortems.

Dev Deletes Entire Production Database, Chaos Ensues

Kevin Fang, 2023. 10-minute explainer.

A step-by-step retelling of the 2017 GitLab outage, in which several backup and replication mechanisms that looked healthy turned out to be empty or broken when a restore was needed.

Read

Release It!: Design and Deploy Production-Ready Software

Michael T. Nygard, 2018, 2nd edition.

Catalogs production failure modes (cascading failures, blocked threads, unbounded result sets) and the stability patterns that keep one broken dependency from taking down the whole product.

Database Reliability Engineering: Designing and Operating Resilient Database Systems

Laine Campbell and Charity Majors, 2017.

Covers backup strategy, restore testing, recovery objectives and replication from the operator's side, including why a backup nobody has restored cannot be trusted.

Primary sources

  • Reference

    Postmortem of database outage of January 31 (GitLab)

    GitLab's own postmortem, with the timeline, the backup and replication methods that failed silently, and the systemic fixes that followed, which also makes it a model for writing one.

  • Reference

    Implementing health checks (Amazon Builders' Library)

    Explains the trade-off between shallow liveness checks and deep dependency checks, and how to keep a health check from reporting healthy while the service can't reach its database, or from turning one dependency failure into a fleet-wide outage.