Neural Language-Model Baselines and Deep-Learning Foundations
machine-learning
language-models
optimization
Derive softmax and cross-entropy, train baseline models, and connect the reference gradients to the reusable PyTorch path.
A Transformer is easier to debug after its objective and gradients have been reduced to inspectable cases. This chapter derives softmax cross-entropy and keeps NumPy implementations as transparent counter-examples. The reusable baseline path then moves to PyTorch tensors and modules, where the same next-token objective can be optimized, inspected, and handed to the decoder in Chapter 04.
The point of the NumPy cells is not to maintain a second training stack. They expose the mathematics and failure modes; proof_lm owns the PyTorch baselines consumed by later chapters.
Next-token likelihood
For a batch of B examples with vocabulary size V, let z_i\in\mathbb{R}^V be the logits and y_i the target index. The softmax distribution is
Teacher forcing supplies the observed previous tokens to every training example. The model therefore learns the conditional distribution at each position while the target sequence remains known; generation later has to feed back its own sampled tokens.
Before introducing a model, make the numerical behavior of these two functions explicit.
loss: 1.636563
row probability sums: [1. 1. 1. 1.]
Subtracting the largest logit before exponentiating leaves every probability unchanged because the same factor cancels in the numerator and denominator. It prevents overflow when a later model produces a large activation. The loss is finite and the row sums equal one, so these invariants can be tested before any optimizer is involved.
Cross-entropy gradients
For one-hot target vectors e_{y_i}, differentiating the log-softmax gives
A finite-difference check compares this hand derivation with the slope measured by perturbing one parameter at a time. The check is small enough to inspect and strong enough to catch a missing batch factor or a wrong target index.
linear loss: 1.343788
maximum gradient error: 1.7439993893475503e-11
The maximum error is at numerical-difference scale, which validates the complete softmax-to-linear chain for this tiny model. The division by batch size is visible in the agreement; omitting it would produce a gradient exactly three times too large here. This is the manual reference against which the multilayer implementation can be checked conceptually.
Bigram baselines
A bigram model assigns one categorical distribution to each preceding token. Let C_{a,b} count transitions from token a to token b. The maximum-likelihood row is
\hat p(b\mid a)=\frac{C_{a,b}}{\sum_u C_{a,u}},
and the minimum achievable average training loss is the empirical conditional entropy
Train a table of logits with gradient descent and compare its loss with this independently calculated number. The comparison tests the data indexing, softmax, gradient, and update rule together.
initial recorded loss: 1.4161
final loss: 0.787097
empirical bigram entropy: 0.785596
The trained table approaches the entropy of the observed transition distribution because each row can represent its own conditional probabilities. It cannot improve on that entropy on the same empirical distribution: cross-entropy equals entropy plus a nonnegative KL divergence, and the divergence is zero at the maximum-likelihood rows. The remaining gap is finite-step optimization, not a missing hidden state.
Fixed-context baselines
The bigram table sees one previous token. A fixed-context MLP sees C previous tokens, looks up an embedding for each, concatenates the embeddings, and maps them through a nonlinear hidden layer:
The context length is a data-shape contract. A model trained with C=3 cannot silently receive a two-token input without padding or a defined shorter-context rule. The next cell keeps the forward cache explicit because the backward pass needs each intermediate array.
The embedding gradient uses np.add.at because the same token can occur several times in one batch. Ordinary indexed assignment would retain only the last contribution and silently change the derivative. The hidden layer gives the model a nonlinear way to combine positions, while the fixed concatenation still makes the receptive field exactly three tokens.
tiny-batch loss: 1.8444 -> 0.001827
main train loss: 1.5849 -> 0.273499
main validation loss: 0.269394
The tiny-batch run is a debugging test, not a generalization result. A model with enough parameters should drive a handful of fixed examples close to zero; failure points to an indexing, gradient, or update error before hyperparameter comparisons begin. The main run separates the decreasing training loss from validation loss. Their gap is a capacity-and-data diagnostic, not proof that the larger model has learned a useful rule.
Sampling and generation
Training selects the parameters that assign probability to observed targets. Generation selects a token from the resulting distribution. Temperature \tau rescales logits as z/\tau; lower values concentrate probability around the largest logits. Top-k sampling sets all but the k largest logits to zero probability before drawing. Greedy decoding is the limiting case k=1.
Use the same prompt and a named random generator for each comparison so the decoding policy, rather than hidden global state, determines the difference.
The NumPy cells make the objective and gradient visible. The cumulative implementation is now a PyTorch module with the same next-token interface. Run the bigram and fixed-context MLP on a small periodic stream, then compare their decreasing losses before the Transformer adds attention and a larger receptive field.
The three samples use identical learned logits but different support and temperature rules, so their continuations need not agree. A random generator makes stochastic sampling reproducible without making it deterministic across policies. The context experiment reports validation loss for a fixed update budget; increasing context changes both the information available and the number of parameters, so it should be read as a controlled baseline rather than an isolated claim about context length.
Summary
Stable softmax and cross-entropy turn next-token likelihood into testable NumPy functions.
The hand-derived linear gradient matches finite differences, including the batch normalization factor.
A bigram table trained by gradient descent approaches the empirical conditional entropy, which is an independent target for the training loop.
A fixed-context MLP adds embeddings and nonlinear capacity; deliberate tiny-batch overfitting detects implementation bugs.
Temperature and top-k modify the sampling distribution after training, while validation loss measures a separate held-out property.
Chapter 04 replaces the fixed local context with causal self-attention and checks the shape and gradient contracts of a decoder block.
Exercises
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P3.1] Softmax gradient
Gradient derivation. Starting from mean softmax cross-entropy, derive the derivative with respect to one logit vector. Explain why the batch-size factor must appear in both the weight and bias gradients.
[P3.2] Sampling policy
Sampling policy. Given logits [2.0, 1.0, 0.0], describe the support of greedy decoding and top-k sampling with k=2. Explain how increasing temperature changes the relative probabilities without changing their ordering.
[P3.3]
PyTorch baseline contract. Explain why the bigram baseline should be evaluated with the previous-token tensor context_inputs[:, -1], while the fixed-context MLP receives the full (batch, context_length) tensor. State one shape assertion that would catch swapping those inputs.
def baseline_input_shapes(context_inputs, bigram_inputs, context_length):# Return a pair of booleans for the bigram and MLP contracts.pass