Layer L1
ML / DL Basics
The layer
What this layer solves
You have trained a network with your own hands, and you can explain why it learned, why it overfits, and why the metrics are chosen the way they are. Everything in L2 uses this layer's vocabulary: the pretraining loss, the splits in fine-tuning, the metrics in evaluation. Without it, LLM training code reads like a pile of API calls. Three conditions for the finish line: train a small network from scratch, explain overfitting and generalization, and say why the evaluation metrics are chosen the way they are.
How the nodes connect into one line
This layer is a line unrolled from the training loop:
- Classical machine learning: learning from samples: the vocabulary. Train / validation / test splits, loss, model families, metrics. Google's crash course covers these with interactive modules.
- Loss and backpropagation: the middle two steps of the loop. Why the loss function is cross-entropy or MSE, and why backprop is just the chain rule run in reverse mode on a computation graph. Write your own autograd by following micrograd and this layer clicks.
- PyTorch: train a network yourself: swap the previous step for a framework. Tensors, autograd, the training loop. After this, no training code looks strange.
- Generalization: explains the curves. Why training and validation loss diverge, how to trade off bias and variance, what regularization does, and the double descent specific to deep networks.
- Classic architectures: CNN, RNN / LSTM, residuals: what the three inventions before the Transformer each solved. CNN's weight sharing, LSTM's gating, the residual's gradient path. The aim is to see what the Transformer kept and what it threw away.
Prerequisite order: classical-ml → backprop-and-loss → pytorch → generalization / classic-architectures. backprop-and-loss connects directly to L0's optimization and information theory.
What a senior engineer can cut
- You don't need all of classical machine learning. Skip SVM, naive Bayes and clustering. Keep linear models, trees and ensembles, and metrics.
- Don't study PyTorch as a framework. Treat it as "numpy that does autograd". Writing the training loop by hand once is enough.
- You don't need to tune CNN and LSTM. You only need to say what inductive bias each one has. Leave the vision-task details until you really use them.
- Only skim Goodfellow's book ch. 5, and use it as a dictionary.
- Skip Keras, TensorFlow and any high-level training library. What they hide is exactly what this layer wants you to see.
Top picks
In reading order; each pick links to its node below.
- 1
Machine Learning Crash Course — Google free
Run through the vocabulary of loss, splits and metrics first → Classical machine learning: learning from examples
- 2
The spelled-out intro to neural networks and backpropagation: building micrograd free
Write autograd in a hundred-odd lines and backprop becomes yours → Loss and backpropagation
- 3
Calculus on Computational Graphs: Backpropagation free
Half an hour on why it is reverse mode and not forward mode → Loss and backpropagation
- 4
Neural Networks: Zero to Hero — Andrej Karpathy free
After micrograd, keep following along, all the way to GPT → PyTorch: train a network yourself
- 5
Learn the Basics — PyTorch tutorials free
The framework's own words on tensors, autograd and the training loop. Use it as a manual → PyTorch: train a network yourself
- 6
An Introduction to Statistical Learning with Applications in Python free
The standard answer on bias–variance and cross-validation → Generalization
- 7
Deep Double Descent: Where Bigger Models and More Data Hurt free
Why large models don't overfit the textbook way. One more layer for interview follow-ups → Generalization
- 8
What CNN, ResNet, RNN and LSTM each fixed in the previous generation, with code → Classic architectures: CNN, RNN / LSTM, residuals
- 9
Understanding LSTM Networks free
Four gate diagrams, the shortest path to seeing what attention replaced → Classic architectures: CNN, RNN / LSTM, residuals
The nodes
Classical machine learning: learning from examples
Half-life: stable
The idea existed before language models, just at a smaller size: a model learns from examples by lowering its loss, a number that says how wrong it is. The traps are the same at any size: test on data it was trained on and it looks better than it is; train too long and it memorizes the examples instead of learning from them (overfitting). You can learn all this on a small tabular problem, where a training run takes seconds. It is the vocabulary for everything after: the loss of an LLM, the metrics of evals, the data splits of fine-tuning all use these words.
- Builds on
- Math: learn it when you need it
- Leads to
- Generalization · PyTorch: train a network yourself
- On the routes
- Foundations (Machine learning, the classical way)
course — The fundamentals — loss, generalisation, splits, classification metrics — in short interactive modules, updated by Google in 2024.
Hands-on project: A baseline you can defend — Train logistic regression and gradient-boosted trees on one tabular dataset with a proper train, validation and test split, and explain one case where accuracy is the wrong metric.
Not picked (3)
- An Introduction to Statistical Learning (ISLP) — Very good, but the bias–variance and cross-validation chapters fit better in the generalization node, and the model-family part overlaps with Google's course.
- Raschka, Machine Learning with PyTorch and Scikit — Learn — paid and thick; the hands-on part is already covered by Google's course and PyTorch: train a network yourself.
- scikit — learn User Guide — 查 API 用,不是学习路径。
Loss and backpropagation
Half-life: stable
A training loop has four steps: the forward pass computes the output, the loss function measures the gap, backpropagation computes the gradients, and the optimizer updates the parameters. This node covers the middle two: why the loss is cross-entropy or MSE (which distribution each one assumes), and why backpropagation is just the chain rule run in reverse mode on a computation graph. It connects the gradients from Optimization and the cross-entropy from Information theory into code that runs. It is everything behind the loss.backward() line in PyTorch: train a network yourself. The reproduction project for L1 is attached here: train a small network yourself and plot the training/validation curves. Generalization then explains why the curves diverge.
- Builds on
- Math: learn it when you need it · Optimization · Information Theory
- Leads to
- Classic architectures: CNN, RNN / LSTM, residuals · PyTorch: train a network yourself · Transformer
video · 2.5 h video + 3 h coding · only 0:00–2:25:52 — Write a scalar autograd engine in 100-odd lines, train a small MLP on it. Backprop becomes code you can write and check yourself: this layer's 'train a small network from scratch'.
article · 30 minutes — Frames backprop as forward and reverse mode differentiation on a graph. One page shows why reverse mode is cheap for many inputs and one loss. It answers 'why not forward mode?'
Hands-on project: Train a small network from scratch and plot training/validation curves — Write a two-layer MLP using only tensor operations (no nn.Module, no optimizer). Train it on MNIST or make_moons, record training and validation loss each epoch, and plot them as one chart. First let it overfit (the curves diverge), then add weight decay or dropout to close the gap. The output is two charts and a short explanation.
Not picked (3)
- 3Blue1Brown「Backpropagation calculus」 — Fourth episode of the series, same source as the intro video in Math: fill in when needed; watching episode one leads into it naturally.
- CS231n notes「Backpropagation, Intuitions」 — Well explained, but Olah's piece is shorter and also covers the forward/reverse mode comparison.
- Goodfellow 等《Deep Learning》ch. 6.5 — Formal but dense; writing micrograd along with it gives more.
PyTorch: train a network yourself
Half-life: stableWho asks for it
PyTorch is how most models get written and trained. The core is a short loop: a batch of data goes through the model, you compute the loss, let autograd work out how each weight should move, and take a step. Write that loop yourself, then build up from a tiny network to a small GPT, and the architecture is no longer just a diagram. It is also the prerequisite for writing GPT from scratch and for fine-tuning.
- Builds on
- Classical machine learning: learning from examples · Math: learn it when you need it · Loss and backpropagation
- Leads to
- Classic architectures: CNN, RNN / LSTM, residuals · Fine-tuning: teach it your task · Transformer
- On the routes
- Foundations (Machine learning, the classical way) · AI Engineer (Train and post-train)
course — Backprop to a GPT in PyTorch, one video at a time — the fastest way for an engineer to read and debug model code.
docs — Tensors, autograd, datasets and the training loop in the framework's own words; a reference you will return to.
Hands-on project: Train a small GPT and change one thing — Reproduce a small model with nanochat or nanoGPT on a modest dataset, change one design choice, and show the ablation with its loss curves.
Not picked (3)
- fast.ai Practical Deep Learning for Coders — Top-down: you get results first with a high-level library. This layer needs you to see the training loop itself, so Zero to Hero fits better.
- Dive into Deep Learning — Already a candidate under Classic architectures: CNN, RNN / LSTM, residuals, so not repeated here.
- PyTorch 官方「60 Minute Blitz」 — Overlaps with Learn the Basics, which is more recent.
Generalization
Half-life: stable
A low loss on the training set does not mean the model is useful. Generalization is about the gap between training error and test error: where it comes from, how to measure it, and how to shrink it. Core vocabulary: bias–variance, capacity, regularization (weight decay, dropout, early stopping, data augmentation), cross-validation, and double descent, which is specific to deep networks. It is the theory behind splitting datasets in Classical machine learning: learning from samples, and the key to reading the curves you plot in the Loss and backpropagation project. By Pretraining and scaling, scaling laws are still, at heart, about how data size, parameter count and generalization error relate.
book · only ch. 2–5 — Ch. 2 draws the bias–variance trade-off and train/test error in one picture; ch. 5 shows why cross-validation estimates generalization error. The standard answer on overfitting.
paper · 1 hour — The classic U-shaped curve falls again in overparameterized deep networks. When asked why huge models don't overfit, you can answer one layer deeper than the textbook.
Hands-on project: Make overfitting happen, then fix it — Take the small network from the [[backprop-and-loss]] project. Raise the parameter count to 10× and then 100×, plot training and validation error against parameter count, and see whether you can reproduce double descent on a small dataset. Write a paragraph explaining the cause of each part of the curve.
Not picked (3)
- Goodfellow 等《Deep Learning》ch. 5, 7 — Thorough, but the two ISLP chapters are shorter and come with code; ch. 5 is already a candidate in Classical machine learning: learning from samples.
- Zhang 等「Understanding deep learning requires rethinking generalization」(2017) — A landmark paper that posed the problem, but the double descent paper gives a more usable picture.
- Bishop, Pattern Recognition and Machine Learning ch. 1 — A solid Bayesian view, but too slow to read.
Classic architectures: CNN, RNN / LSTM, residuals
Half-life: stable
Three key inventions came before the Transformer, and each solved one specific problem. CNN uses local connections and weight sharing to build the translation invariance of images into the structure, cutting the parameter count from the millions of a fully connected network to something trainable. RNN lets a network handle variable-length sequences, and the gates in LSTM fix its vanishing gradients and its failure to remember distant information. Residual connections let gradients pass straight through dozens or hundreds of layers. Without them there would be no deep networks, and no skip around every sublayer in the Transformer. We study these not to use them, but to understand what the Transformer kept (residuals, normalization), what it threw away (recurrence), and why. The vision encoder in a multimodal model is still a convolutional network or a ViT, and this node is its prerequisite.
- Builds on
- Loss and backpropagation · PyTorch: train a network yourself
- Leads to
- Multimodal
book · only ch. 7–10 — Four chapters: CNN, modern CNN (ResNet), RNN, modern RNN (LSTM). Each names the flaw of the last generation, then gives runnable PyTorch. After it you can say what each one solved, in one line.
article · 30 minutes — Four gate diagrams show why LSTM remembers distant information and a plain RNN cannot. It is the shortest path to understanding what the Transformer's attention replaced.
- CS231n Convolutional Neural Networks for Visual Recognition — course notes: Convolutional Networks free
course · 1 hour — Convolution, stride, padding, receptive field, parameter sharing, each with a countable diagram. You can compute any layer's output size and parameters, the base skill for vision backbones.
Hands-on project: Run one task on three architectures — On a small sequence classification dataset, train an MLP, a 1D CNN and an LSTM. Record parameter count, training time and validation accuracy. Write a paragraph on what each architecture's inductive bias corresponds to in the data.
Not picked (3)
- He 等「Deep Residual Learning for Image Recognition」(2015) — The original residual paper, but d2l ch. 8 already explains it and gives an implementation.
- Goodfellow 等《Deep Learning》ch. 9 — 10 — The content is right, but it has no code and no residuals.
- fast.ai Practical Deep Learning for Coders — Top-down, using high-level libraries, so you can't see the architecture itself.
Self-check
- Train a two-layer MLP to a reasonable validation accuracy using only tensor operations, with no nn.Module and no optimizer.
- Plot the training / validation loss curves, point out the epoch where overfitting starts, and close the gap with one regularization method.
- Explain, for one scenario, why accuracy is the wrong metric, what to use instead, and why.
- Say in three sentences each what CNN, LSTM and residual connections solve.
- State how the bias–variance tradeoff relates to double descent.
Hands-on projectTrain a small network from scratch and plot the training / validation curves: make it overfit first, then fix it. Produce two plots and a short explanation (hang it on Loss and backpropagation).
Translated from the author's Chinese notes by a model; the Chinese page is the original.