Building Software

Engineering Fundamentals for the Agent Era

Contents Section 7, Operations

Observability

Mistakes to catch in review

  1. Whole request bodies logged, including passwords, tokens and personal data.

  2. Errors logged without a request ID, inputs or cause, so they can't be traced or reproduced.

  3. Alerts on CPU and memory while user-facing errors climb unnoticed.

  4. A user ID used as a metric label, creating millions of time series and a large monitoring bill.

Logs, metrics and traces that let you answer new questions about a running system and learn it is broken before users tell you.

Topics

Structured Logging
Machine-readable log events with consistent fields, useful context and no secrets.
Metrics
Counters, gauges and histograms, and the RED and USE methods for choosing what to measure.
Distributed Tracing
Following one request across services with trace and span IDs to see where time and errors go.
SLIs, SLOs and Error Budgets
Defining reliability from the user's side and deciding how much failure is acceptable.
Symptom-Based Alerting
Paging people when users are hurting, and keeping every alert actionable.

You understand it when you can

  • Define a service level indicator and objective for a service, and the alert that fires when the objective is at risk.
  • Follow one request across two services using logs, traces and a correlation ID.
  • Add structured logging to a function so a future failure can be diagnosed from the log alone.

Drill

An agent added logging to a login endpoint that writes the full request body at info level, emits a login metric labeled with each user's ID, and alerts when CPU exceeds 80 percent. Find the data leak, the metrics-cost problem and the outage this alerting would miss.

Start here

Watch

The RED Method: How To Instrument Your Services

Tom Wilkie, 2018. 20-minute talk.

The creator of the RED method explains why every service should export Rate, Errors and Duration, and how RED differs from Gregg's USE method for resources, which gives a concrete answer to what a service should measure.

Read

Observability Engineering: Achieving Production Excellence

Charity Majors, Liz Fong-Jones and 2 others, 2026, 2nd edition.

A near-complete rewrite that covers structured events, tracing, SLO-based alerting and telemetry cost, and adds material on validating AI-generated code in production.

Implementing Service Level Objectives: A Practical Guide to SLIs, SLOs, and Error Budgets

Alex Hidalgo, 2020.

Walks through choosing SLIs from the user's point of view, setting targets, and using error budgets and burn rates to decide when to page and when to slow down.

Learning OpenTelemetry: Setting Up and Operating a Modern Observability System

Ted Young and Austin Parker, 2024.

Explains how OpenTelemetry ties traces, metrics and logs together through shared context propagation, and how to instrument services and run the Collector in practice.

Primary sources

  • Reference

    Alerting on SLOs (The Site Reliability Workbook, Chapter 5)

    Works through six alerting strategies step by step, ending at multiwindow, multi-burn-rate alerts that page when the error budget is burning, where CPU thresholds only fire on symptoms users may never feel.

  • Manual

    Prometheus: Metric and label naming

    The official warning that every unique label combination is a new time series and that labels must never hold unbounded values such as user IDs, which is the metrics-cost trap in this subsection's drill.