The Agent Harness: Building a Coding Agent from Scratch

agents
engineering
reliability
Build and measure a headless async coding-agent harness from streaming transport to offline evaluation.

A language model can produce the next message in a conversation, but it cannot by itself read a repository, change a file, or know what happened after an action. The harness supplies those missing channels: it preserves state, translates model proposals into typed tool calls, executes permitted actions, returns observations, and records enough evidence to evaluate the result.

This course builds a headless async coding agent in the Python package projects/agent-harness/. The path starts with a deliberately small think–act–observe loop, then adds streaming transport, typed tools, repository operations, context management, instructions and memory, permissions, sandboxing, hooks, sub-agents, MCP, and an evaluation scorecard. It is for readers who want to understand the mechanisms between a model endpoint and a coding system they can inspect and constrain.

The target is a measurable coding-agent harness, not a reproduction of a branded product or a general model-training course. The core experiments use deterministic fakes and disposable workspaces, so the protocol and its failure modes remain visible even without an API key.

What you will finish with

The completed course leaves you with:

  • a reusable Python package containing an OpenAI-compatible streaming client, event model, typed tool registry, file and shell tools, session and agent loop, context controls, configuration and memory, safety and sandbox backends, hooks, sub-agent delegation, and an MCP bridge;
  • regression tests and offline fixtures that exercise each boundary without contacting a model provider; and
  • a capstone scorecard that judges scripted trajectories against task, tool, safety, context, and evidence criteria.

The scope is intentionally bounded. The package has no user interface or command-line entry point, makes no claim of production-grade isolation, and does not train or fine-tune a language model. Each chapter isolates one mechanism, shows what it guarantees, records what remains outside that guarantee, and adds an observable contract plus a mechanism that measures or constrains the new failure modes introduced by each layer.

The learning path

Each chapter adds an artifact to the package and a falsifiable experiment or contract check1. Read 00. Architecture and evidence contract before Chapter 01; it defines the shared vocabulary and the evidence standard used throughout.

Phase Chapter Build and evidence
Orientation 00. Architecture and evidence contract Map the harness, execution paths, trajectory contract, and capstone acceptance conditions.
Part I: The Engine Exposed 01. Model versus harness Separate a stateless completion from a think–act–observe loop and validate episode structure.
02. Streaming client Assemble streamed text, tool calls, usage, retries, and terminal errors without silent loss.
03. Tool protocol Make tool names, arguments, results, and errors typed and inspectable.
Part II: Making It a Coding Agent 04. Coding tools Read and mutate a disposable repository through exact file and shell contracts.
05. Agent loop Close the session state transition from model turn to tool observation and final answer.
Part III: Controlling It 06. Context management Bound history with pruning, compaction, token accounting, and continuation state.
07. Instructions and memory Load layered instructions and persistent memory without hiding their provenance.
08. Permissions and sandboxing Separate approval decisions from host and Docker execution boundaries.
09. Hooks Attach lifecycle checks and observability while containing hook failures.
Part IV: Scaling Out, Then Proving It 10. Sub-agents and patterns Delegate bounded work with explicit roles, budgets, and isolated or shared context.
11. Model Context Protocol Adapt external tool schemas and transports at a typed boundary.
12. Capstone Integrate the real package on offline fixtures and produce an evidence-backed scorecard.

Prerequisites

Tooling:

  • Python 3.14+ and comfort with functions, classes, async/await, and virtual environments.
  • uv or an equivalent way to install the repository environment, plus pytest for the package tests.
  • A shell and Git. The later chapters operate on disposable repositories and subprocesses.
  • Basic HTTP and JSON knowledge. Server-Sent Events and tool-call schemas are built from first principles in Chapters 02 and 03.

Assumed knowledge:

  • message roles and the idea that a model endpoint maps conversation history to a response;
  • ordinary Python testing and error handling; and
  • enough systems judgment to distinguish a model proposal, a tool observation, and an evaluator’s verdict.

Execution at a glance

The default path is local and offline: run the notebooks with the repository environment, use deterministic fake clients, and exercise tools in .tmp/ fixtures. No GPU, API key, or paid service is required for the complete course path.

A live-provider path is optional. Set one of the package’s supported API-key and base-URL environment variables to connect the streaming client to an OpenAI-compatible endpoint; this path introduces network dependence and usage cost, so it is not the source of the deterministic evidence. Docker is likewise optional and appears as one sandbox backend in Chapter 08.

Important notes

The notebooks are the argument and the backing package is the reusable implementation. Later chapters import the package rather than copying earlier cells, so a change to a boundary can be tested across the full harness. The package’s own tests and the capstone’s scripted model keep the main evidence path offline.

Keep credentials out of notebooks and Git. The live client reads configuration from environment variables, while the course’s standard fixtures use fake clients. A passing trajectory is not automatically a correct trajectory: the harness records events and observations, and the judge applies explicit task criteria to produce the scorecard.

Back to top

Footnotes

  1. Every layer that adds agency must also add a mechanism that measures and constrains its new failure modes. A layer that only adds capability is cannot be considered finished.↩︎