Tool calling is a posttraining outcome, not a prompt-engineering feature bolted onto a finished model. The model must choose among available tools, emit arguments that satisfy a typed schema, decide when answering directly is better than calling, read results back, and continue within a step budget. Each of those is a separate learned behavior with a separate failure mode, and the evaluation protocol keeps them separate.
Evaluation protocol
On held-out requests, report exact tool choice and argument correctness, schema-validity rate, when-to-call accuracy on prompts where calling helps and prompts where it does not, hallucinated-tool and hallucinated-parameter rates, recovery success after a deliberately corrupted tool result, multi-turn completion within the step budget, and the calibration curve of call decisions. Hold out tools and their schemas entirely from training, not just their names: generalization to an unseen schema is a different claim from generalization to an unseen phrasing of a seen tool.
Summary
Tool training is SFT over scripted trajectories with call-span masks, optionally sharpened by DPO pairs on borderline prompts (correct call versus hallucinated tool, well-scoped call versus unnecessary call). The evaluation protocol, not the training loss, defines what was actually learned.
Exercises
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P11.3]
A model emits a known tool name but omits a required argument. Which part of the tool-call protocol fails, and what structured result should the training loop record?
Back to top