Build a composable harness for perplexity, calibration, memorization, overlap, and generation diversity.
A lower next-token loss is evidence about one prediction task, not a complete description of a language model. This chapter builds the evaluation harness that later posttraining chapters can reuse. The NumPy functions keep calibration, canary, overlap, and diversity definitions inspectable; the shared ProofLM evaluators attach explicit checkpoint-stage identities to masked likelihood and verifier reports.
Each metric is a small function with explicit inputs and no hidden model state. That design makes it possible to rerun the same suite after a checkpoint, a data change, or a posttraining update and to report the measurements together rather than selecting a single favorable score.
Perplexity
For target tokens y_i and logits z_i, the token negative log-likelihood is
and perplexity is \exp(\operatorname{NLL}). Aggregate the numerator and token count before exponentiating. Averaging sequence-level perplexities would give short sequences the same weight as long ones and would not equal the corpus-level likelihood.
The evaluator below accepts records with split, source, and token IDs, plus a predictor that returns one logit vector for each input token. The model is fitted only on training records so the grouping can expose both source and split effects.
The grouped report keeps token counts beside NLL and perplexity, so a reader can see how much evidence supports each number. Training-like rows should usually be easier for a fitted predictor than shifted rows, but that difference is not itself a capability measure. Source grouping can reveal that an aggregate validation value is being driven by one easy source while another source has much higher uncertainty.
The predictor is a closure over a fixed probability table, while the evaluator knows nothing about how that table was made. Later chapters can replace the predictor with a Transformer checkpoint without changing the grouping logic.
Calibration
For a next-token prediction, let c_i be the maximum predicted probability and let r_i indicate whether the selected token is correct. Partition confidences into bins I_b. The expected calibration error is
A reliability table must retain count, mean confidence, and empirical accuracy for every nonempty bin. ECE compresses those discrepancies to one number, so the table remains the primary diagnostic.
The table distinguishes two kinds of calibration failure. A model can be overconfident when its accuracy is below its mean confidence, or underconfident when accuracy is higher. ECE weights each discrepancy by the number of predictions in its bin, but it does not identify which examples caused the discrepancy. Keep the bin rows and the prediction-level data when calibration affects a later call-or-no-call decision.
Memorization canaries
Insert a unique string into the training corpus and evaluate the conditional probability of its continuation from a prefix. Compare a model fitted with the canary against the same model fitted without it, holding the alphabet and smoothing constant. The probe is deliberately local: a high recovery probability shows that the model retained the sequence, not that it can generalize the surrounding task.
A character bigram is sufficient to make the exposure effect measurable. Repeating the canary increases the counts of its internal transitions, while the held-out control shares the general format but not the exact transition.
def fit_character_bigram(texts, alphabet, smoothing=0.05): symbol_to_id = {symbol: index for index, symbol inenumerate(alphabet)} counts = np.full((len(alphabet), len(alphabet)), smoothing, dtype=float)for text in texts: tokens = np.asarray([symbol_to_id[symbol] for symbol in text], dtype=np.int64) np.add.at(counts, (tokens[:-1], tokens[1:]), 1.0)return counts / counts.sum(axis=1, keepdims=True), symbol_to_iddef continuation_logprob(probabilities, prefix, continuation):ifnot prefix:raiseValueError("prefix must contain one symbol") previous = prefix[-1] total =0.0for token in continuation: total += np.log(probabilities[previous, token]) previous = tokenreturnfloat(total)def greedy_continuation(probabilities, prefix, length): generated =list(prefix)for _ inrange(length): generated.append(int(np.argmax(probabilities[generated[-1]])))return generatedbase_text ="the model reads data and reports loss. "*6canary ="<Q7|zebra|X9>"control ="<Q7|zebra|Y9>"alphabet =sorted(set(base_text + canary + control))canary_ids = np.asarray([alphabet.index(symbol) for symbol in canary], dtype=np.int64)control_ids = np.asarray([alphabet.index(symbol) for symbol in control], dtype=np.int64)without_canary, symbol_to_id = fit_character_bigram([base_text], alphabet)with_canary, _ = fit_character_bigram([base_text, canary], alphabet)prefix = canary_ids[:-2].tolist()continuation = canary_ids[-2:].tolist()seen_score = continuation_logprob(with_canary, prefix, continuation)unseen_score = continuation_logprob(without_canary, prefix, continuation)recovered = greedy_continuation(with_canary, prefix, len(continuation))print("canary continuation log-prob, seen:", round(seen_score, 4))print("canary continuation log-prob, unseen:", round(unseen_score, 4))print("greedy recovery:", recovered == canary_ids.tolist())print("control differs at final transition:", control_ids[-3:].tolist())assert seen_score > unseen_scoreassert recovered == canary_ids.tolist()print("exposure by repetition count")exposure_scores = []for repetitions in (1, 3, 8): model, _ = fit_character_bigram([base_text] + [canary] * repetitions, alphabet) score = continuation_logprob(model, prefix, continuation) exposure_scores.append(score)print(f"repetitions={repetitions}: log-prob={score:.4f}")assert exposure_scores[0] < exposure_scores[-1]
canary continuation log-prob, seen: -1.4793
canary continuation log-prob, unseen: -6.3561
greedy recovery: True
control differs at final transition: [8, 3, 5]
exposure by repetition count
repetitions=1: log-prob=-1.4793
repetitions=3: log-prob=-0.6399
repetitions=8: log-prob=-0.2671
The seen model assigns more probability to the canary continuation and greedily recovers it, while the model without the inserted sequence has only the smoothing floor for the rare final transitions. The repetition sweep makes exposure a graded quantity rather than a binary label. This result is a memorization measurement under a specified model and prefix; it does not establish that every occurrence of the string would be recoverable from every prompt.
N-gram overlap
For a generated sequence g and a set of training sequences D, define the n-gram overlap rate as
Use unique candidate n-grams in the numerator so a repeated phrase is not counted several times merely because it appears repeatedly in the same generation. Report several values of n: unigram overlap mostly reflects vocabulary, while longer overlap is more sensitive to copied local sequences.
def ngram_set(sequence, n): sequence =tuple(sequence)if n <1:raiseValueError("n must be positive")return { sequence[index:index + n]for index inrange(max(0, len(sequence) - n +1)) }def ngram_overlap(generated, training_sequences, n=3): candidate = ngram_set(generated, n)ifnot candidate:return0.0 training =set().union(*(ngram_set(sequence, n) for sequence in training_sequences))returnlen(candidate & training) /len(candidate)def overlap_report(generations, training_sequences, orders=(1, 2, 3)):return { n: float(np.mean([ ngram_overlap(generation, training_sequences, n)for generation in generations ]))for n in orders }training_sequences = ["the cat sat on the mat".split(),"the dog ran to the park".split(),"a model predicts the next token".split(),]generations = ["the cat sat on the mat".split(),"the model predicts a token".split(),"new birds fly above water".split(),]overlap = overlap_report(generations, training_sequences)for order, value in overlap.items():print(f"mean unique {order}-gram overlap: {value:.3f}")assert ngram_overlap(training_sequences[0], training_sequences, n=3) ==1.0assert ngram_overlap(generations[-1], training_sequences, n=3) ==0.0assertset(overlap) == {1, 2, 3}
mean unique 1-gram overlap: 0.667
mean unique 2-gram overlap: 0.417
mean unique 3-gram overlap: 0.333
The overlap report preserves the order of the comparison. A generation can have high unigram overlap because it uses the training vocabulary while having low trigram overlap, or it can copy a long phrase while changing individual word frequencies. No order alone distinguishes useful reuse from memorization, so interpret the rates with canary probes and held-out loss.
Generation diversity
A generation set can be diverse because its samples differ, because its token inventory is broad, or because it avoids repeating short phrases. Measure these separately:
distinct_n is the fraction of unique n-grams among all generated n-grams.
unique_sequence_ratio is the fraction of samples that are distinct as whole sequences.
mean_pairwise_jaccard compares n-gram sets between pairs and detects collections that reuse the same local phrases.
These metrics describe output variation. They do not establish correctness, factuality, or calibration.
def distinct_n(generations, n): all_ngrams = [ngram for generation in generations for ngram in ngram_set(generation, n)] total =sum(max(0, len(generation) - n +1) for generation in generations)returnlen(set(all_ngrams)) / total if total else0.0def unique_sequence_ratio(generations):returnlen({tuple(generation) for generation in generations}) /len(generations)def mean_pairwise_jaccard(generations, n=2): sets = [ngram_set(generation, n) for generation in generations] scores = []for left inrange(len(sets)):for right inrange(left +1, len(sets)): union = sets[left] | sets[right] scores.append(len(sets[left] & sets[right]) /len(union) if union else1.0)returnfloat(np.mean(scores)) if scores else0.0def generation_diversity(generations):return {"distinct-1": distinct_n(generations, 1),"distinct-2": distinct_n(generations, 2),"unique-sequence-ratio": unique_sequence_ratio(generations),"mean-pairwise-2gram-jaccard": mean_pairwise_jaccard(generations, n=2), }repetitive_generations = ["the cat sat on the mat".split(),"the cat sat on the mat".split(),"the cat sat near the mat".split(),]varied_generations = ["red birds fly above water".split(),"small models read heldout data".split(),"green fish swim below ice".split(),]repetitive_metrics = generation_diversity(repetitive_generations)varied_metrics = generation_diversity(varied_generations)print("repetitive:", {key: round(value, 3) for key, value in repetitive_metrics.items()})print("varied: ", {key: round(value, 3) for key, value in varied_metrics.items()})assert repetitive_metrics["unique-sequence-ratio"] < varied_metrics["unique-sequence-ratio"]assert varied_metrics["distinct-2"] > repetitive_metrics["distinct-2"]
The repetitive set has duplicate whole sequences and shared local phrases, so its unique-sequence ratio and distinct-2 score are lower. The varied set scores better on those dimensions, but a diverse set could still be uniformly incorrect. Keep diversity as a companion to likelihood, calibration, and task-specific checks rather than using it as a quality proxy.
Evaluation suite
A reusable harness returns a structured report whose fields can be compared across checkpoints. The composition function below does not average unlike quantities. It keeps grouped perplexity, the calibration table and ECE, memorization evidence, overlap by order, and diversity metrics in separate namespaces. Later stages can add instruction accuracy or tool-call validity without changing these primitive functions.
The pure metric functions above define the measurements. The package-backed evaluators add the artifact boundary: masked likelihood reports an evaluator identity, while proof validity is scored by the independent verifier rather than by the model’s own likelihood. These compact reports are the interface consumed by later posttraining chapters.
The final object is intentionally a report of vectors, tables, and rates rather than a leaderboard scalar. A later checkpoint can be evaluated with the same records, prompts, and generation settings, then compared field by field. That comparison makes a regression visible when validation perplexity improves but calibration, canary recovery, or copied n-grams worsen.
Summary
Perplexity is computed from aggregate token NLL and is grouped by split and source so data volume and domain effects remain visible.
Reliability bins pair mean confidence with empirical correctness; ECE summarizes but does not replace the bin table.
A canary inserted into training provides a controlled conditional-recovery probe whose exposure can be varied by repetition count.
N-gram overlap measures local reuse at several orders, while diversity metrics separate token variety, whole-sequence variety, and shared phrases.
The evaluation suite is composed from small pure functions and returns separate namespaces, allowing later posttraining stages to rerun the same measurements.
Chapters 08 through 12 add posttraining objectives and tool-use behavior; each stage should rerun this base suite and add its own task-specific checks.
Exercises
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P7.1] Expected calibration error
Calibration calculation. Three predictions have confidences [0.2, 0.4, 0.9] and correctness [1, 0, 0]. Using three bins with one prediction in each, compute the ECE and identify whether the high-confidence prediction is overconfident.
[P7.2] Generation memorization audit
Generation audit. Explain why a high distinct-2 score does not prove correctness. Then state how n-gram overlap and a canary probe provide different evidence about memorization.
[P7.3]
Evaluator identity. Explain why likelihood and proof validity need separate evaluator IDs even when they are reported for the same checkpoint. Give one example of a regression that a lower perplexity would fail to reveal.
def evaluator_regression(lower_perplexity, proof_validity):# Return whether the checkpoint passes both independent gates.pass