7
Operations: Running Software in Reality
Seeing what production is doing and changing it without betting the system on each release.
Why it matters when agents do the typing
Production is where software meets real users, real data and real failures, and someone has to answer when it breaks. Debugging and observability tell you what is happening and why, redundancy lets the system survive what nobody predicted, and safe change lets a team ship many times a day.
-
Core 13
Depth: do
Debugging
Finding the cause of a failure by reproducing it and running experiments, instead of guessing at fixes.
A mistake it teaches you to catch: A symptom patched with a null check, a retry or a broad try/catch while the actual cause stays in place.
-
Core 14
Depth: do
Observability
Logs, metrics and traces that let you answer new questions about a running system and learn it is broken before users tell you.
A mistake it teaches you to catch: Whole request bodies logged, including passwords, tokens and personal data.
-
Depth: do
Reliability, Redundancy and Recovery
Designing systems that survive component failures, recovering data and service when they don't, and learning from every incident.
A mistake it teaches you to catch: Backups that were never restored and turn out to be empty, partial or corrupt when they are needed.
-
Core 15
Depth: do
Shipping Change Safely
Version control, deployment pipelines, migrations, feature flags and rollback: the machinery for changing a live system many times a day.
A mistake it teaches you to catch: A migration that locks a large table, or drops a column the running version still reads.
-
Depth: do
Performance and Capacity
Measuring where time and money go, finding bottlenecks before users do, and planning for load.
A mistake it teaches you to catch: Code optimized where it is not the bottleneck, while the real cost sits in a query or a network call.