The Anatomy of Machine Learning: Data, Models, Loss, and Optimization

How a model learns from examples: data, a family of functions, a loss, an optimizer, and the held-out test that separates learning from memorization.

Sylla N'falyJan 05, 20266 min read

A program that converts Celsius to Fahrenheit is written by hand: someone knows the rule, , and types it. Machine learning starts from the other end. We have pairs , measurements with some noise, and no rule. Learning means choosing, from a family of candidate functions, the one whose predictions stay closest to , and hoping it also works on pairs we have not seen yet.

Everything below is one piece of that sentence: the pairs (data), the family (model), "closest" (loss), "choosing" (optimization), and "pairs we have not seen" (generalization).


The setup

We assume the examples come from an unknown distribution . We never see that distribution, only a finite dataset drawn from it.

The model defines a hypothesis space : all the functions it can represent as its parameters vary. What we want is the with the smallest error on new samples from . We cannot measure that directly, so we minimize the average error on and check the result on data kept aside.

Four components

01 / The Data

A finite sample from the unknown distribution .

02 / The Model

The family of functions , indexed by the parameters .

03 / The Loss Function

A single number that measures how wrong the predictions are.

04 / The Optimizer

The procedure that searches for good parameters; for neural networks, a gradient-based update.

Forward pass and loss evaluation against the targets
Input x goes through the model with parameters θ to a prediction ŷ, which the loss function compares with the true label y to produce a single loss value.

Model capacity and inductive bias

The model decides which shapes of function can be learned at all. A linear model can only represent straight lines (hyperplanes in higher dimensions): a strong inductive bias, an assumption built in before seeing any data. A deep neural network can represent far more complex functions, and its biases come from its architecture: a convolutional network, for instance, assumes that nearby pixels are related.

How flexible the family is, is the model's capacity.

The bias-variance tradeoff

  • Too little capacity (high bias): the model cannot capture the structure in the data. It underfits: errors are high on the training set and on new data.
  • Too much unchecked capacity (high variance): the model fits the noise along with the signal. It overfits: low training error, high error on new data.
Underfitting, a good fit, and overfitting
Three fits of the same points: a straight line that misses the trend, a smooth curve that follows it, and a jagged line that passes through every point.

Measuring error: the loss

The loss turns "how wrong" into a single number . Two common choices:

  • Regression, mean squared error:
  • Classification over K classes, cross-entropy: , where is the one-hot label and the predicted probabilities. Only the probability given to the correct class counts, and the loss grows quickly as that probability approaches 0.

Training looks for the parameters that make the loss smallest:


The training loop

For neural networks, the minimum is found step by step: compute the gradient of the loss, the direction in which it increases fastest, and move the parameters a little in the opposite direction. How long each step should be is a question of its own, covered in the article on gradient descent and step size (in French).

The full training loop
Loop: batch of inputs, forward pass, loss, backward pass computing gradients, parameter update, then the next batch.
  1. Forward pass

    Pass a batch of inputs through the model to get predictions .

  2. Loss computation

    Compare the predictions with the targets through the loss .

  3. Backward pass (backpropagation)

    Compute the gradient of the loss with respect to every parameter, using the chain rule: .

  4. Update

    Move the parameters against the gradient, scaled by the learning rate :

Gradient descent on a two-parameter loss surfaceJacopo Bertolotti, Wikimedia Commons, CC0
Animated trajectory of gradient descent moving downhill across a loss surface toward its minimum.

A minimal implementation in PyTorch

The Celsius idea, with a different rule: data generated from plus noise, and a linear model that has to find the two numbers back.

# train_regression.py
import torch
import torch.nn as nn
import torch.optim as optim
 
# 1. Synthetic data: y = 2x + 1, plus noise with standard deviation 0.1
torch.manual_seed(42)
X = torch.randn(100, 1)
y = 2.0 * X + 1.0 + 0.1 * torch.randn(100, 1)
 
# 2. Model: one weight and one bias
model = nn.Linear(in_features=1, out_features=1)
 
# 3. Loss and optimizer
criterion = nn.MSELoss()
optimizer = optim.SGD(model.parameters(), lr=0.1)
 
# 4. Training loop
for epoch in range(101):
    y_pred = model(X)
    loss = criterion(y_pred, y)
 
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
 
    if epoch % 25 == 0:
        print(f"Epoch {epoch:03d} | Loss: {loss.item():.4f}")
 
w, b = model.weight.item(), model.bias.item()
print(f"Learned: y = {w:.2f}x + {b:.2f}")

Each step here uses all 100 examples, so optim.SGD is doing plain (full-batch) gradient descent; the "stochastic" part appears when each step sees only a random batch.

Two things about the output can be predicted without running it. The noise has a standard deviation of 0.1, so no straight line can reach a mean squared error much below on this data: the printed loss should fall and level off near that value. And the learned line should be close to . The exact printed numbers depend on the PyTorch version and the random initialization, so they are not reproduced here.


Generalization and the held-out split

An example split: 70% training, 15% validation, 15% test
A dataset bar split into training (70%, used by gradient descent), validation (15%, model selection and early stopping) and test (15%, final evaluation, untouched until the end).

What matters is performance on data the model has never seen. The training set is used to fit the parameters, the validation set to choose between models and settings (capacity, learning rate, when to stop), and the test set is used once, at the end, to estimate how the chosen model will do on new data. The 70/15/15 proportions are a convention, not a rule; with little data, cross-validation reuses it more efficiently.


References

  1. Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2017). Understanding Deep Learning Requires Rethinking Generalization. ICLR 2017.
  2. Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., & Tang, P. T. P. (2017). On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. ICLR 2017.
  3. Dinh, L., Pascanu, R., Bengio, S., & Bengio, Y. (2017). Sharp Minima Can Generalize For Deep Nets. arXiv:1703.04933.
All posts