Watch
"Performance Matters" by Emery Berger
Shows how layout effects and noise mislead naive benchmarks, and introduces causal profiling (Coz) to find the code whose speedup would actually make the program faster.
Engineering Fundamentals for the Agent Era
Contents Section 7, Operations
Code optimized where it is not the bottleneck, while the real cost sits in a query or a network call.
A cache with no invalidation that serves stale data, or a cache key that serves one user's data to another.
A service that works for ten users and falls over at a thousand because nobody considered the connection pool or a rate limit.
Autoscaling or model usage with no ceiling, turning a traffic spike into a very large bill.
Measuring where time and money go, finding bottlenecks before users do, and planning for load.
An agent sped up a slow product page by memoizing a pricing function, and added a cache keyed only on the product ID for a response that includes the viewer's saved cart. Profile to find where the page actually spends its time, and find the cache bug that shows one user another user's cart.
Watch
Shows how layout effects and noise mislead naive benchmarks, and introduces causal profiling (Coz) to find the code whose speedup would actually make the program faster.
Explains why averages and even p99 hide what users experience, and how coordinated omission makes most load-testing tools under-report tail latency.
Shows real sites whose caches stored one user's authenticated page and served it to others because the cache key and rules ignored who was asking, the per-user caching bug in this subsection's drill.
The standard reference for performance methodology (USE method, workload characterization) and for profiling CPU, memory, disk and network before changing any code.
Its opening chapters define load, throughput, latency percentiles and tail-latency amplification, and later chapters explain how caches, replicas and derived data go stale.
Shows how to measure real capacity ceilings from production data, forecast growth, and plan resources and autoscaling limits before traffic arrives.
Paper
Shows mathematically why rare slow responses dominate user-facing latency once a request fans out to many servers, and describes hedged requests and other ways to reduce the tail.
RFC
The HTTP caching standard, defining cache keys, Vary, Cache-Control private and no-store, and freshness, the rules that decide whether a shared cache may store and reuse a per-user response.