Carry one model through every pipeline stage and characterize exactly where reliable tool behavior breaks.
Carry one small decoder-only model through the complete pipeline: pretrain, instruction-tune, apply at least one preference or RL method, train tool calling, and evaluate every stage against the same executable specification. The deliverable is a reproducible study with fixed budgets, multiple seeds, and documented failures, in that order of emphasis.
Comparison protocol
Evaluate five models under identical conditions: the base pretrained model, the instruction-fine-tuned model, the preference-optimized model, the tool-trained model, and a non-learning baseline (template-matched tool selection). Freeze architecture, tokenizer, prompts, decoding settings, target specification, and compute accounting across all five. Run at least three seeds for early experiments and five for final comparisons where compute allows.
Evaluation suite
For every model: in-distribution tool-use tests, held-out instances, held-out tools and schemas, paraphrased requests, multi-turn episodes, shifted contexts, malformed-result recovery, correction requests, repeated sampling for stochastic failure rates, and general capability checks with tool scaffolding removed. The research questions the suite answers:
Which pipeline stage contributes most to held-out tool-calling performance?
Does the model follow the tool-use specification or superficial formatting cues? (Paraphrase and held-out-schema results.)
Does it exploit evaluator weaknesses where they exist? (Weak-checker versus ground-truth scores.)
Does it respond to correction after a failed call? (Correction test results.)
Which failures are visible in outputs alone, and which need internal inspection?
Baselines
Before trusting any comparison: freeze model, data, and dependency versions; verify split independence with hash overlap; test the target evaluator on hand-constructed examples; confirm held-out tools never appeared in training or prompt construction; verify a deliberately incorrect output scores lower; then freeze checkpoints and run the full suite, repeated with independent seeds. Perform causal diagnosis only after the behavioral baseline is stable, or the intervention targets noise.
Deliverables and completion criteria
Submit the versioned target specification; the data, training, and evaluation harness; base, SFT, preference/RL, and tool-trained checkpoints; held-out and shifted test sets; compliance, calibration, and capability metrics with means and standard deviations over seeds; a failure taxonomy with concrete examples; at least one causal diagnosis; and a report of limitations and reproducibility details. The capstone is complete when another implementer can reproduce the evaluation, verify the target evaluator, obtain comparable results across seeds, and identify at least one concrete limitation of the claims.
Current capstone evidence boundary
The end-to-end record below connects the implemented contracts and deterministic proof evaluator, but it does not pretend that a CPU smoke run is a standard-scale result. A complete capstone still needs multiple seeded checkpoints, held-out tool/schema splits, and the CUDA or other hardware qualification described in the plan.
import torchfrom proof_lm.logic import generate_examplesfrom proof_lm.posttraining import verifier_rewardfrom proof_lm.tools import ProofToolRegistrycapstone_examples = generate_examples(count=2, seed=14)[:2]capstone_record = {"stages": ["pretrain_smoke", "sft_contract", "preference_contract", "tool_contract"],"proof_rewards": [verifier_reward(example.proof_text) for example in capstone_examples],"tool_registry_id": ProofToolRegistry().registry_id,"cuda_available": torch.cuda.is_available(),"seed": 14,"scale_gate": "standard 1B run not claimed by this CPU smoke",}print("capstone stages:", " -> ".join(capstone_record["stages"]))print("proof reward checks:", capstone_record["proof_rewards"])print("tool registry:", capstone_record["tool_registry_id"])print("CUDA available:", capstone_record["cuda_available"])print("scale gate:", capstone_record["scale_gate"])assert capstone_record["proof_rewards"] == [1.0, 1.0]
capstone stages: pretrain_smoke -> sft_contract -> preference_contract -> tool_contract
proof reward checks: [1.0, 1.0]
tool registry: tools-0d40b519fe44f7c4
CUDA available: False
scale gate: standard 1B run not claimed by this CPU smoke
import torchfrom proof_lm.logic import generate_examplesfrom proof_lm.posttraining import verifier_rewardfrom proof_lm.tools import ProofToolRegistrycapstone_examples = generate_examples(count=2, seed=14)capstone_record = {"stages": ["pretrain_smoke", "sft_contract", "preference_contract", "tool_contract"],"proof_rewards": [verifier_reward(example.proof_text) for example in capstone_examples],"tool_registry_id": ProofToolRegistry().registry_id,"cuda_available": torch.cuda.is_available(),"seed": 14,"scale_gate": "standard 1B run not claimed by this CPU smoke",}print("capstone stages:", " -> ".join(capstone_record["stages"]))print("proof reward checks:", capstone_record["proof_rewards"])print("tool registry:", capstone_record["tool_registry_id"])print("CUDA available:", capstone_record["cuda_available"])print("scale gate:", capstone_record["scale_gate"])assert capstone_record["proof_rewards"] == [1.0, 1.0]
Exercises
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P14.1] Reproducible capstone records
Name four records required for a reproducible capstone comparison and explain why one final score is insufficient.
[P14.2] Threats ruled out by records
Name four records that must accompany a capstone comparison for it to be reproducible, and give one distinct threat each record rules out.
[P14.3]
Why does the current capstone record keep a scale gate instead of reporting the CPU smoke as a final model result?