Building Software

Engineering Fundamentals for the Agent Era

Contents Section 4, Math and Algorithms

Probability and Statistics

Mistakes to catch in review

  1. Average latency reported while a slow tail, which most users hit at least once per session, stays hidden.

  2. A prompt or model change declared an improvement from a handful of runs, with no estimate of run-to-run variance.

  3. Failures treated as independent when they share a cause, such as two replicas in the same data center.

  4. A flaky test called fixed after a single green run.

Distributions, sampling, variance and significance: the tools for judging whether a measurement, an experiment or an eval result means anything.

Topics

Probability Basics
Events, independence, conditional probability and expected value.
Distributions and Percentiles
Normal and long-tailed distributions, and why percentiles describe user experience better than averages.
Sampling and Variance
How much a measurement moves between runs, what a confidence interval says, and how many samples are enough.
Experiments and Significance
A/B tests, hypothesis testing and the common ways experiments fool the people who run them.
Base Rates and False Positives
Why rare events make detectors, alerts and classifiers noisy, and how Bayes' rule corrects intuition.

You understand it when you can

  • Compute the mean, median and 99th percentile of a latency sample and explain which one users feel.
  • Decide how many runs you need before trusting a difference between two versions of a nondeterministic system.
  • Explain why a 99 percent accurate detector can still produce mostly false alarms when the thing it detects is rare.

Drill

An agent reports a flaky test as fixed after three consecutive green runs; before the fix it failed about one run in ten. Compute the chance of three greens with the bug still present, and say how many consecutive green runs it takes before a bug that is still there would produce that streak less than 1 percent of the time.

Start here

Watch

The medical test paradox, and redesigning Bayes' rule

Grant Sanderson, 2020. 21-minute explainer.

Works through why a highly accurate test for a rare condition produces mostly false positives, which carries over directly to flaky-test reruns, alerts and classifiers.

Watch

"How NOT to Measure Latency" by Gil Tene

Gil Tene, 2015. 43-minute talk.

Shows why averages and naive percentiles hide the slow tail that most users hit, and how coordinated omission corrupts latency measurements.

But what is the Central Limit Theorem?

Grant Sanderson, 2023. 31-minute explainer.

Builds intuition for how sample means spread and shrink with sample size, the basis for confidence intervals and for deciding how many runs are enough.

Read

Think Stats: Exploratory Data Analysis

Allen B. Downey, 2025, 3rd edition. Free to read online.

Teaches distributions, percentiles, estimation and hypothesis testing through runnable notebooks, so the ideas turn into code a developer can apply to their own measurements.

Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing

Ron Kohavi, Diane Tang and Ya Xu, 2020.

Covers the ways experiments mislead people in practice, including underpowered tests, peeking and invalid randomization, drawn from running A/B tests at Microsoft, Google and LinkedIn.

Introduction to Probability

Joseph K. Blitzstein and Jessica Hwang, 2019, 2nd edition.

The Harvard Stat 110 text on conditional probability, independence, expectation and distributions, for readers who want the foundations behind the rules of thumb.

Primary sources