Failure Modes, Specification, and Causal Debugging

reliability
interpretability
Make the training target executable, train against a deliberately weak evaluator, and diagnose the gap with behavioral tests before interventions.

A model can raise its training score while missing the intended objective, and the divergence is not rare or mysterious: it follows from optimizing a measurable proxy. This chapter makes that divergence the subject. The working method is to define a target that is executable enough for repeatable evaluation, train against a weakened version of it, and measure the gap with a ground-truth evaluator.

Target behavior

Define the target S with exact tests, a simulator with a known reward, a procedural evaluator, a structured output schema, or a task-specific verifier. Version the specification and test it on hand-constructed examples before trusting any model score against it: a verifier that disagrees with hand-labeled examples measures the verifier, not the model.

Failure modes

Mode Mechanism Detection
Specification error The written objective does not express the intended behavior Hand-labeled examples disagree with the evaluator
Proxy optimization An easier measurable signal correlates with the objective in training data Gap between training-like and held-out or shifted compliance
Reward hacking The policy exploits a weakness in the evaluator Held-out ground-truth score flat while training reward climbs
Goal misgeneralization Behavior that fit the objective in training pursues something else in new situations G_{\mathrm{gap}} = A_{\mathrm{train}} - A_{\mathrm{shift}} grows
Distribution shift Test inputs differ from anything in training Shifted-context suite
Correction failure Feedback after a mistake does not change the next attempt Correction test: same task, feedback, remeasure

Behavioral probes

Build a small synthetic environment where the intended objective is known and the evaluator can be weakened on purpose: a sequence rule that changes mid-episode, a task with a resource constraint, a structured answer with a factual invariant, tools usable only when a precondition holds. Before intervening inside the model, interrogate it from the outside:

  • Paraphrase sweeps: does compliance survive rewording?
  • Composition: do known skills combine, or does only the trained combination work?
  • Repeated sampling: separate stochastic failure (varies across samples) from systematic failure (stable across samples).
  • Correction tests: after a failed attempt and explicit feedback, does the next attempt change?

Causal interventions

When behavior cannot localize the cause, intervene: train a probe to read a property from activations, then follow it with ablation (zero or baseline a unit, head, or slice) and activation patching (run clean and corrupted inputs, splice a selected activation slice from the clean run into the corrupted run, and measure restoration). Probe accuracy shows information is present, not that the model uses it; only the patch or ablation effect establishes use. The output of the experiment is a causal effect estimate on the target behavior plus an unrelated-behavior control, repeated on a held-out checkpoint.

Metrics

Report objective-compliance rate in-distribution and held-out, distribution-shift compliance, the goal-generalization gap G_{\mathrm{gap}}, evaluator-versus-ground-truth disagreement, hacking rate against the weak evaluator, correction success rate, capability retained, and calibration and abstention quality. Sparse feature dictionaries and activation steering are research extensions outside this course.

Summary

The behavioral toolkit (paraphrase, composition, repeated sampling, correction) decides what is failing; the causal toolkit (probe, ablation, patch) decides where. Run them in that order, and keep the evaluator versioned throughout.

Executable proxy-gap check

A weak evaluator is useful only when its disagreement with the target is made visible. The shared proof verifier supplies the ground-truth column, while the intentionally substring-based checker supplies the proxy column. The shortcut case passes the proxy and fails the target, which is the smallest reproducible reward-hacking experiment.

from proof_lm.logic import generate_examples
from proof_lm.posttraining import verifier_reward

ground_truth = generate_examples(count=1, seed=13)[0].proof_text
weak_checker = lambda text: float("verified" in text.lower())
cases = [
    ("correct proof", ground_truth),
    ("shortcut output", "verified, but not a proof"),
]
scorecard = [
    (name, weak_checker(text), verifier_reward(text))
    for name, text in cases
]
print("case / weak reward / ground truth:", scorecard)
assert scorecard[0][2] == 1.0
assert scorecard[1][1] == 1.0 and scorecard[1][2] == 0.0
case / weak reward / ground truth: [('correct proof', 0.0, 1.0), ('shortcut output', 1.0, 0.0)]

Exercises

Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.

[P13.1] Probe versus causal use

A probe predicts whether an activation came from a shifted example. Does that prove the model uses shift information? Explain the next experiment.

[P13.2] Shortcut evaluator

Construct a checker that can be gamed by a shortcut, and specify one held-out test that separates proxy compliance from objective compliance.

[P13.3]

In the proxy-gap check, why is the shortcut output evidence about the evaluator rather than about model capability? Name the paired metric that exposes it.

Back to top