Building Software

Engineering Fundamentals for the Agent Era

Contents

9

Directing Agents: Delegation, Review and Accountability

Why it matters when agents do the typing

Whoever approves a change answers for it, whether a person or an agent typed it, which makes the agents part of the system you are responsible for. Directing them well means knowing how they fail and how they are attacked, giving them clear intent, limiting what they can touch, and reviewing what comes back. This section comes last because it draws on everything before it: you can't review what you don't understand.

  1. Core 19

    How Agents Work and How They Fail

    The mechanics of language models and agent loops, and the failure modes that follow from them: invented APIs, agreement with a bad spec, confident wrong fixes, injected instructions and different results on every run.

    A mistake it teaches you to catch: A call to a client-library option that does not exist, such as a retries setting the library never had, written with full confidence.

    Depth: do
  2. Core 20

    Prompting and Specifying Work for Agents

    The acceptance criteria from Intent, plus the three things a prompt adds: the context the agent needs, the boundaries of what it may touch, and the check that proves it is done.

    A mistake it teaches you to catch: An agent told to 'make the tests pass' that deletes or weakens the tests.

    Depth: do
  3. Core 21

    Reading and Reviewing Code You Did Not Write

    Reading diffs and unfamiliar code critically: what to check first, where agent changes typically go wrong, and how to review more code than you could ever write.

    A mistake it teaches you to catch: A change that passes the tests but alters behavior outside the requested task.

    Depth: do
  4. Evals: Testing AI Behavior

    Measuring nondeterministic systems with datasets, graders and statistics, so a prompt or model change is judged on evidence.

    A mistake it teaches you to catch: A prompt change shipped because it looked better on three examples.

    Depth: explain
  5. Orchestration, Permissions and Guardrails

    Running agents in loops and pipelines with limited permissions, budgets and approval points, so their mistakes stay small.

    A mistake it teaches you to catch: A CI agent given the organization-wide deploy token because scoping one to the repository was more setup, so any repository it touches can ship to production.

    Depth: explain
  6. Core 22

    Accountability: Owning Code You Did Not Type

    What stays yours when agents do the work: the approval, the incident, the license and the dependencies, plus the skills to act on them yourself.

    A mistake it teaches you to catch: A 600-line agent change approved within a minute, with approval treated as a formality.

    Depth: do