Capstone: Proving the Harness with Offline Evidence
agents
evaluation
reliability
capstone
The capstone integrates the mechanisms built across the course: an async agent loop, typed tools, context and session state, safety policy, hooks, sub-agent boundaries, and an MCP bridge. Integration is not a victory lap. It is the point at which a local guarantee can fail because another layer ignored it.
The chapter therefore makes one modest claim falsifiable: the offline harness satisfies a stated set of contracts on deterministic fixtures, and each contract can be made to fail by removing its corresponding control. No model call, MCP server, subprocess, or secret is required. The evidence is narrower than a production benchmark, but it is repeatable and inspectable. See the sub-agent patterns and MCP bridge chapters for the two external-authority boundaries being integrated here.
Evaluation contract
The capstone artifact is not just the final assistant string. A passing run must leave observable evidence for four questions:
Task: did the scripted agent complete the requested repository change?
Authority: did tools, sub-agents, and external schemas retain their boundaries?
Control: did safety, context, and loop controls catch the planted failures?
Evidence: can a judge score those claims from events, files, and structured observations?
The offline runner below uses the real package classes wherever a boundary exists. Fakes appear only at the model and MCP transport edges, where deterministic inputs are more informative than a flaky live service.
Workspace and registry
Use .tmp/ rather than the operating system temporary directory so the fixture’s lifecycle is visible to the repository workflow and remains gitignored. The registry is the same default registry used by Session, including the delegation tool introduced in Chapter 10.
The fixture is now attached to a normal Config and ToolRegistry. YOLO is explicit here because the current Agent loop dispatches through the registry but does not itself call an approval callback. The safety policy is still evaluated separately below; naming that limitation prevents the integration demo from overstating what the current loop enforces.
Scripted agent run
The fake client emits the same StreamEvent objects that LLMClient produces. On turn one it requests write_file; on turn two it returns a final answer. The agent, session, registry dispatch, file tool, event stream, and stored usage are real. Only the model decision is fixed.
class ScriptedClient:def__init__(self, turns: list[list[StreamEvent]]):self.turns = turnsself.calls =0asyncdef chat_completion(self, messages, tools=None, stream=True): events =self.turns[self.calls]self.calls +=1for event in events:yield eventscripted = ScriptedClient( [ [ StreamEvent(type=StreamEventType.TEXT_DELTA, text_delta=TextDelta("I will create the report. "), ), StreamEvent(type=StreamEventType.TOOL_CALL_COMPLETE, tool_call=ToolCall( call_id="write-1", name="write_file", arguments={"path": "report.txt", "content": "verified offline\n"}, ), ), StreamEvent(type=StreamEventType.MESSAGE_COMPLETE, usage=TokenUsage(prompt_tokens=42, completion_tokens=12, total_tokens=54), ), ], [ StreamEvent(type=StreamEventType.TEXT_DELTA, text_delta=TextDelta("Report created and verified."), ), StreamEvent(type=StreamEventType.MESSAGE_COMPLETE, usage=TokenUsage(prompt_tokens=58, completion_tokens=7, total_tokens=65), ), ], ])session = Session(config, client=scripted, registry=registry)agent = Agent(config, session=session)asyncdef collect_events():return [event asyncfor event in agent.run("Create report.txt with the fixture result.")]events =await collect_events()event_types = [event.type.value for event in events]print(event_types)print("final text:", [event.data["content"] for event in events if event.type== AgentEventType.TEXT_COMPLETE])assert scripted.calls ==2assert (fixture_root /"report.txt").read_text(encoding="utf-8") =="verified offline\n"assert AgentEventType.TOOL_CALL_COMPLETE.value in event_typesassert event_types[-1] == AgentEventType.AGENT_END.value
['agent_start', 'text_delta', 'text_complete', 'tool_call_start', 'tool_call_complete', 'text_delta', 'text_complete', 'agent_end']
final text: ['I will create the report. ', 'Report created and verified.']
The successful artifact is stronger than a printed “done”: the file content, two model calls, a tool-complete event, and a terminal agent event all agree. The session accumulated 54 + 65 = 119 total tokens across turns. The event stream is an observability surface that later evaluation code can score without re-running the model.
Journal replay
The session keeps an append-only journal of user messages, assistant tool calls, tool results, turns, and usage. Replay provides a second route to the same state. This is useful in evaluation because a stored trajectory can be inspected or re-judged without contacting the model again.
Replay establishes state equivalence, not semantic correctness. A journal can faithfully preserve a bad tool call. That is why the suite below scores safety, authority boundaries, and bounded termination in addition to task completion.
Held-out fixtures
A suite is useful only when its cases are fixed before an ablation is run. These seven cases cover the integrated loop and one boundary from each hardening layer. They are small enough to understand individually, but they exercise files, events, schemas, async dispatch, safety classification, hook validation, pruning, and loop detection through the project APIs.
@dataclass(frozen=True)class EvalCase: case_id: str family: str description: str dimensions: tuple[str, ...]@dataclass(frozen=True)class FixtureObservation: case_id: str checks: dict[str, bool] details: dict[str, object]SUITE_CASES = ( EvalCase("agent-loop","integration","complete a scripted write and expose its lifecycle events", ("task_success", "observability"), ), EvalCase("path-boundary","safety","reject a file path outside the workspace and classify a dangerous shell command", ("safety",), ), EvalCase("hook-contract","extension","reject an ambiguous lifecycle hook before external code can run", ("hooks",), ), EvalCase("delegation-boundary","authority","keep an investigator read-scoped and bounded", ("delegation",), ), EvalCase("mcp-bridge","protocol","translate and invoke a namespaced external tool offline", ("mcp",), ), EvalCase("context-budget","control","prune old context while retaining the system prompt and tail", ("context",), ), EvalCase("loop-stop","control","detect a repeated action before it consumes unbounded turns", ("bounded_termination",), ),)print([(case.case_id, case.dimensions) for case in SUITE_CASES])assertlen(SUITE_CASES) ==7assertlen({case.case_id for case in SUITE_CASES}) ==len(SUITE_CASES)
Each case names the dimension it owns. This prevents a single aggregate number from hiding which mechanism failed. The fixture suite is held out from the model because it does not ask a model to discover the expected answer; it asks the harness to preserve explicit invariants.
Boundary fixtures
The runner below is longer than a unit test because it records details alongside booleans. A failed row should explain what was observed, not merely lower a score. The variant argument removes one named control for ablation; the full path always uses the actual project mechanism.
The first three runners cover the integrated loop, permissions, and delegation. The remaining runners exercise hook configuration, MCP schema namespacing and dispatch, context pruning, and loop detection. Each runner accepts the same variant string, so removing one control changes only the observation owned by that control.
The MCP check requires both a successful call and a namespaced schema, so transport success cannot hide a collision-prone registration. The context check separately asserts preservation of the system message and recent tail. The hook case stays offline by testing HookConfig validation rather than launching user code; Chapter 9 already established the subprocess timeout contract.
These are narrow mechanism checks. They do not claim that a namespaced MCP description is free of prompt injection, that mechanical pruning preserves every fact, or that rejecting an ambiguous hook proves a future veto protocol correct.
Ablation matrix
Convert each fixture observation into the project evaluator’s EvalTask, Trajectory, and JudgeResult types. The conversion deliberately keeps task success and contract compliance separate: a failed boundary produces a failed trajectory even when the Python runner itself completed. Every variant writes its trajectories under a distinct directory, then the Scorecard aggregates the same fixed task suite.
The full harness passes all seven fixtures. Every one-control ablation fails exactly one owned case, while the other six remain green. This is the required causal sanity check: the evaluator can observe a control disappearing, and the failure is attributed to the matching dimension instead of reducing an unexplained aggregate score.
The score difference is intentionally coarse. With two equally weighted judge criteria, an ablated case contributes zero while every unaffected case contributes one. The useful evidence is the per-case failure and stored observation, not the third decimal place in the mean.
Reproducibility manifest
A result without an implementation identifier is not replayable. This offline run has no provider dashboard or decoding parameters, so record those fields explicitly as not applicable. Hash the harness source and effective configuration, version the evaluator protocol, and point to the suite and scorecards written above.
The offline result proves that the current package composes on deterministic fixtures: the agent loop can perform a real file write and replay its journal; path and approval checks reject the planted hazards; hook configuration rejects an ambiguous contract; delegation preserves a read-only tool boundary and turn budget; MCP translation preserves a namespace through dispatch; context pruning protects the system message and recent tail; and the loop detector catches a repeated action. The persisted suite, trajectories, scorecards, and manifest make each observation inspectable without contacting a model.
It does not prove the broader production claim from the course plan. A held-out repository benchmark still needs development and held-out tasks with no shared files, at least three seeds per configuration, fixed live-model decoding parameters, wall-clock and provider cost reconciliation, calibrated LLM-judge prompts, sandbox escape attempts, and failure exemplars from trajectories longer than 50 calls. Those requirements cannot be manufactured by an offline notebook.
Treat this chapter as the deterministic base layer for that study. Clone a repository fixture per attempt, freeze suite.json before the first held-out run, store one trajectory per seed and variant, and keep the offline matrix as a preflight test. If an external evaluator cannot reproduce these seven rows first, a larger benchmark score is not trustworthy.