Failure Modes, Specification, and Causal Debugging
reliability
interpretability
Make the training target executable, train against a deliberately weak evaluator, and diagnose the gap with behavioral tests before interventions.
A model can raise its training score while missing the intended objective, and the divergence is not rare or mysterious: it follows from optimizing a measurable proxy. This chapter makes that divergence the subject. The working method is to define a target that is executable enough for repeatable evaluation, train against a weakened version of it, and measure the gap with a ground-truth evaluator.
Target behavior
Define the target S with exact tests, a simulator with a known reward, a procedural evaluator, a structured output schema, or a task-specific verifier. Version the specification and test it on hand-constructed examples before trusting any model score against it: a verifier that disagrees with hand-labeled examples measures the verifier, not the model.
Failure modes
Mode
Mechanism
Detection
Specification error
The written objective does not express the intended behavior
Hand-labeled examples disagree with the evaluator
Proxy optimization
An easier measurable signal correlates with the objective in training data
Gap between training-like and held-out or shifted compliance
Reward hacking
The policy exploits a weakness in the evaluator
Held-out ground-truth score flat while training reward climbs
Goal misgeneralization
Behavior that fit the objective in training pursues something else in new situations
Feedback after a mistake does not change the next attempt
Correction test: same task, feedback, remeasure
Behavioral probes
Build a small synthetic environment where the intended objective is known and the evaluator can be weakened on purpose: a sequence rule that changes mid-episode, a task with a resource constraint, a structured answer with a factual invariant, tools usable only when a precondition holds. Before intervening inside the model, interrogate it from the outside:
Paraphrase sweeps: does compliance survive rewording?
Composition: do known skills combine, or does only the trained combination work?
Repeated sampling: separate stochastic failure (varies across samples) from systematic failure (stable across samples).
Correction tests: after a failed attempt and explicit feedback, does the next attempt change?
Causal interventions
When behavior cannot localize the cause, intervene: train a probe to read a property from activations, then follow it with ablation (zero or baseline a unit, head, or slice) and activation patching (run clean and corrupted inputs, splice a selected activation slice from the clean run into the corrupted run, and measure restoration). Probe accuracy shows information is present, not that the model uses it; only the patch or ablation effect establishes use. The output of the experiment is a causal effect estimate on the target behavior plus an unrelated-behavior control, repeated on a held-out checkpoint.
Metrics
Report objective-compliance rate in-distribution and held-out, distribution-shift compliance, the goal-generalization gap G_{\mathrm{gap}}, evaluator-versus-ground-truth disagreement, hacking rate against the weak evaluator, correction success rate, capability retained, and calibration and abstention quality. Sparse feature dictionaries and activation steering are research extensions outside this course.
Summary
The behavioral toolkit (paraphrase, composition, repeated sampling, correction) decides what is failing; the causal toolkit (probe, ablation, patch) decides where. Run them in that order, and keep the evaluator versioned throughout.
Executable proxy-gap check
A weak evaluator is useful only when its disagreement with the target is made visible. The shared proof verifier supplies the ground-truth column, while the intentionally substring-based checker supplies the proxy column. The shortcut case passes the proxy and fails the target, which is the smallest reproducible reward-hacking experiment.
from proof_lm.logic import generate_examplesfrom proof_lm.posttraining import verifier_rewardground_truth = generate_examples(count=1, seed=13)[0].proof_textweak_checker =lambda text: float("verified"in text.lower())cases = [ ("correct proof", ground_truth), ("shortcut output", "verified, but not a proof"),]scorecard = [ (name, weak_checker(text), verifier_reward(text))for name, text in cases]print("case / weak reward / ground truth:", scorecard)assert scorecard[0][2] ==1.0assert scorecard[1][1] ==1.0and scorecard[1][2] ==0.0
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P13.1] Probe versus causal use
A probe predicts whether an activation came from a shifted example. Does that prove the model uses shift information? Explain the next experiment.
[P13.2] Shortcut evaluator
Construct a checker that can be gamed by a shortcut, and specify one held-out test that separates proxy compliance from objective compliance.
[P13.3]
In the proxy-gap check, why is the shortcut output evidence about the evaluator rather than about model capability? Name the paired metric that exposes it.