Logs, metrics and traces that let you answer new questions about a running system and learn it is broken before users tell you.
Topics
- Structured Logging
- Machine-readable log events with consistent fields, useful context and no secrets.
- Metrics
- Counters, gauges and histograms, and the RED and USE methods for choosing what to measure.
- Distributed Tracing
- Following one request across services with trace and span IDs to see where time and errors go.
- SLIs, SLOs and Error Budgets
- Defining reliability from the user's side and deciding how much failure is acceptable.
- Symptom-Based Alerting
- Paging people when users are hurting, and keeping every alert actionable.