Building Software

Engineering Fundamentals for the Agent Era

Contents Section 9, Directing Agents

How Agents Work and How They Fail

Mistakes to catch in review

  1. A call to a client-library option that does not exist, such as a retries setting the library never had, written with full confidence.

  2. An agent that agrees a 5-second timeout in the spec is fine for a report that takes two minutes to build, because it goes along with what it was told.

  3. A fix that makes the failing test pass by special-casing the test's input, reported as solving the bug.

  4. An agent that follows an instruction planted in a fetched web page or tool result, such as a line telling it to update the deploy credentials.

The mechanics of language models and agent loops, and the failure modes that follow from them: invented APIs, agreement with a bad spec, confident wrong fixes, injected instructions and different results on every run.

Topics

Next-Token Prediction and Hallucinated APIs
Models generate the most plausible continuation, which is why they invent functions, flags and packages that look right and do not exist.
Preference Tuning and Sycophancy
Training on human approval rewards agreeable answers, so models tend to go along with a flawed spec or a wrong claim.
Context Windows and Dropped Constraints
What a model can see at once, and why instructions far back in a long session or buried in a large context get dropped.
Sampling and Nondeterminism Between Runs
Why the same request can produce different code on different runs, and what that means for reproducing and reviewing agent work.
Agent Loops and Confident Wrong Fixes
How an agent that iterates until a check passes can patch symptoms, weaken tests or report success it never verified.
Prompt Injection Through Tool Results and Fetched Pages
Instructions and data arrive in the same channel, so text in a web page, file, issue or tool output can steer the agent that reads it.

You understand it when you can

  • Explain the mechanism behind a hallucinated API, and check any unfamiliar call against the library's real source or documentation.
  • Run the same agent task twice, compare the outputs, and explain which differences come from sampling and which from changed context.
  • Give an agent a spec with a deliberate flaw and record whether it pushes back or builds the flaw.
  • Explain why an instruction inside a fetched document can steer an agent, and why filtering alone does not prevent it.

Drill

An agent's transcript for 'add retries to the HTTP client' shows four things: it passed a retries option to a client whose documented options are timeout, headers and baseUrl; it agreed that a 5-second timeout is enough for a report that takes two minutes to build; it fetched a documentation page containing a line telling it to disable certificate checks, and did so; and it reports that the tests pass. For each, say whether it is a defect, name the mechanism if it is, and name the evidence that would settle the one you cannot judge from the transcript.

Start here

Watch

[1hr Talk] Intro to Large Language Models

Andrej Karpathy, 2023. 60-minute talk.

Explains an LLM as a next-token predictor that 'dreams' plausible text, which is why it invents APIs, and ends with a security section on jailbreaks and prompt injection through content the model reads.

Watch

Generative AI's Greatest Flaw - Computerphile

Mike Pound, 2025. 12-minute explainer.

Mike Pound compares indirect prompt injection with SQL injection and explains why the fix does not carry over: an LLM has no separate channel that keeps instructions apart from data.

Read

What Is ChatGPT Doing ... and Why Does It Work?

Stephen Wolfram, 2023. Free to read online.

Walks through next-token probabilities and temperature sampling step by step, which covers both why output sounds plausible without being true and why two runs of the same prompt differ.

Build a Large Language Model (From Scratch)

Sebastian Raschka, 2024.

Builds tokenization, attention, pretraining and instruction fine-tuning in code, so the reader sees exactly where the context window and the sampling step come from.

Prompt Engineering for LLMs: The Art and Science of Building Large Language Model-Based Applications

John Berryman and Albert Ziegler, 2024.

Written by two engineers who built GitHub Copilot, it explains how a model reads a document and predicts its continuation, and how that reading causes dropped constraints and confident fabrication.

Primary sources