Skip to main content

Layer L0

Math and CS foundations

The layer

What this layer solves

When I read a paper, I don't skip the formulas. When my training code has a numerical problem, I know where to look. L0 is not a course to "finish". It is a toolbox I come back to when I need it. When I read Transformer and see QKᵀ/√d, I can say what projection this is and why we divide by the square root of d. When I look at a loss curve, I know the number is "how many bits per token on average". When I see Adam's four hyperparameters, I can say what each one controls. The finish line is just two things: derive the gradient of softmax and cross-entropy by hand, and read the formulas of any L2 paper without skipping a line.

How the nodes connect into one line

The entry point is Math: fill in when needed. It splits this layer into four sub-topics, one node each:

  • Linear algebra: what one network layer computes. Matrix multiplication as a geometric action, how shapes line up. attention is three matrix multiplications plus one softmax.
  • Probability and statistics: why the model's output is a distribution, and why the training objective is likelihood. It directly explains the shape of the loss and what temperature means when sampling.
  • Information theory: the negative log of the likelihood is the cross-entropy, the exponent of the cross-entropy is perplexity, and KL divergence is the "don't drift too far from the reference model" constraint in post-training.
  • Optimization: how gradients are computed and turned into parameter updates. Backpropagation is the chain rule, and an optimizer is a second round of processing on the gradient.

The prerequisite links between these four nodes are loose. The real order is "go back and fill in whichever block you can't read in the upper layers". Linear algebra gets used most often. Probability and information theory go as a pair. Optimization follows right after L1's Loss and backpropagation.

What a senior engineer can cut

  • Don't study any math book from the start. MML and Chan's probability book can both be read chapter by chapter. Read only the chapters named in the nodes.
  • In linear algebra, skip determinants and the proofs about abstract vector spaces. You only need intuition for matrix multiplication, transpose, norms and eigendecomposition.
  • In probability, skip measure theory and hypothesis testing. You only need random variables, conditional probability, expectation and maximum likelihood.
  • In optimization, skip convex optimization theory. The loss surface in deep learning is not convex.
  • In information theory, you only need three concepts: entropy, cross-entropy and KL. That is the amount of one article.
  • This layer has no separate node for algorithmic complexity. The part interviews ask for is in L6, and attention's O(n²) is covered in Transformer.

Top picks

In reading order; each pick links to its node below.

  1. 1

    But what is a neural network? — 3Blue1Brown free

    See what a layer computes first, then learn why it computes it → Math: learn it when you need it

  2. 2

    Essence of linear algebra free

    Treat matrix multiplication as a geometric action. Every Transformer diagram after this is built on it → Linear algebra

  3. 3

    Mathematics for Machine Learning — Deisenroth, Faisal and Ong free

    A free textbook you can read chapter by chapter: one chapter each on linear algebra, vector calculus and probability → Math: learn it when you need it

  4. 4

    The Matrix Calculus You Need For Deep Learning free

    Only the slice of matrix calculus that deep learning uses. You need it to derive gradients by hand → Math: learn it when you need it

  5. 5

    Introduction to Probability for Data Science free

    From joint distributions to maximum likelihood, and derive "cross-entropy = negative log-likelihood" → Probability and Statistics

  6. 6

    Visual Information Theory free

    Entropy, cross-entropy and KL drawn out in one hour. After this, perplexity means something → Information Theory

  7. 7

    An overview of gradient descent optimization algorithms free

    From SGD to Adam: what each optimizer changed from the one before → Optimization

The nodes

Math: learn it when you need it

Half-life: stable

layer N⋮layer 1weights × vectorThecatsatbecauseitwastiredAttention: “it” looks back, mostly at “cat”

The model itself is a stack of layers, and a layer is mostly multiplication: a vector times a big table of numbers called weights. A large model has billions of weights. Each layer of a Transformer also lets every token look back at the tokens before it — that is attention, and it is how "it" gets linked to the thing it refers to. Using a model does not need this math. Changing a model does. So wait until a later step needs it: when you can't read a formula, come back and fill in that one piece, rather than finishing a whole book first. Linear algebra lets you see "what one layer computes", calculus lets you see "where the gradient comes from", and probability lets you see "why the output is a distribution".

Leads to
Loss and backpropagation · Classical machine learning: learning from examples · GPU basics: why it's used, and where the bottleneck is · PyTorch: train a network yourself
On the routes
Foundations (How the models work)
  • But what is a neural network? — 3Blue1Brown free

    video · 19 min — The linear algebra of a network seen rather than derived: the picture of what a layer computes that the rest of the map builds on.

  • Mathematics for Machine Learning — Deisenroth, Faisal and Ong free

    book · only ch. 2–6 — Free, and written to be read in parts: linear algebra, vector calculus and probability, each chapter usable alone when a step asks for it.

  • The Matrix Calculus You Need For Deep Learning free

    article · 3 h — Only the slice of matrix calculus deep learning uses — Jacobians and the chain rule over vectors; by the end you can take a two-layer network's gradients by hand, this step's project.

Hands-on project: Backpropagation by hand, then checked — Work out the gradients of a two-layer network for one example on paper, then compare them with PyTorch's autograd.

Not picked (3)
  • Gilbert Strang, MIT 18.06 — A full-semester linear algebra course. Here you only need to see what one layer computes; look it up as needed in the linear algebra node.
  • Goodfellow 等《Deep Learning》ch. 2 — 4 — The math is written densely. Deisenroth's MML is also free and easier to read chapter by chapter.
  • Khan Academy 多元微积分 — Broad coverage but no matrix form; the Parr & Howard piece fills exactly that gap.

Linear algebra

Half-life: stable

linear-algebratransformerno prerequisite

Vectors, matrices, matrix multiplication, transpose, norms, eigendecomposition, SVD. One layer in deep learning is "multiply by a matrix, add a vector, pass through a nonlinearity." Attention is three matrix multiplications plus one softmax. So this layer only asks you to understand what matrix multiplication does geometrically and how shapes line up. You do not need abstract linear spaces. It is the most used of the four sub-topics under 数学:用到时再补: 损失与反向传播 needs it to write the Jacobian, and every diagram in the Transformer is built from it.

Leads to
Transformer
  • Essence of linear algebra free

    video · 16 episodes, about 3 hours — Draws matrix multiplication, change of basis, determinants and eigenvectors as geometric moves. After it, QKᵀ and projections in attention read as pictures, not symbols.

Hands-on project: Write attention by hand in numpy — Take a 4×8 input matrix. Write the Q, K, V projections, QKᵀ/√d, softmax and the weighted sum by hand, noting the shape at each step. Then match the result against torch.nn.functional.scaled_dot_product_attention.

Not picked (3)
  • Gilbert Strang, MIT 18.06 — A full semester course. Too much for an engineer; go back for proofs when needed.
  • Mathematics for Machine Learning ch. 2 — 4 — already the lead pick in 数学:用到时再补, so not repeated here.
  • Linear Algebra Done Right (Axler) — Leans on proofs and skips matrix computation, which does not match how ML uses it.

Probability and Statistics

Half-life: stable

probabilitydecoding-and-inferenceno prerequisite

Random variables, conditional probability, expectation and variance, common distributions, maximum likelihood estimation. A language model's output is a distribution, not an answer. Training maximizes the likelihood of the data, and the sampling temperature changes the shape of that distribution. You can't explain any of this without the vocabulary in this node. It pairs with Information Theory: the negative log of the likelihood is the cross-entropy. Within the four sub-topics of Math: fill in as needed, this one decides whether you can explain why the loss looks the way it does, and what temperature and top-p actually do in Decoding and Inference.

Leads to
Decoding and inference
  • Introduction to Probability for Data Science free

    book · only ch. 2–5 — Free probability text with Python code. Its likelihood chapters show why cross-entropy is negative log-likelihood; by chapter 8 you can derive the softmax + cross-entropy gradient.

Hands-on project: Derive cross-entropy from maximum likelihood — Write the likelihood for a classification task, take the negative log, and show it is cross-entropy. Then differentiate with respect to the softmax input to get p − y, and check it with PyTorch autograd.

Not picked (3)
  • Blitzstein & Hwang, Introduction to Probability (Stat 110) — A classic, but written for statistics majors. Too long and too many exercises for an engineer, and I found no independent endorsement strong enough.
  • Seeing Theory(Brown) — Great visuals, but it stops at intuition. You can't derive the likelihood or the gradient from it.
  • Mathematics for Machine Learning ch. 6 — Already among the picks in the intro of Math: fill in as needed.

Information Theory

Half-life: stable

information-theorybackprop-and-losspretraining-and-scalingno prerequisite

Entropy, cross-entropy, KL divergence, perplexity. It answers three questions you meet every day in engineering: what the loss value actually means (average bits per token), why perplexity is the yardstick for pretraining, and what the "don't drift too far from the reference model" constraint in RLHF and DPO is. It builds on Probability and statistics, and Math: fill in when needed lists it as one of four sub-topics. Upward it connects to Loss and backpropagation (cross-entropy as the loss) and Pretraining and scaling (the vertical axis of scaling laws is this quantity).

Leads to
Loss and backpropagation · Pretraining and scaling
  • Visual Information Theory free

    article · 1 hour — Area charts draw entropy, cross-entropy and KL. You can then say: cross-entropy is the cost of coding with the wrong distribution. Prerequisite for deriving it, perplexity, and DPO's KL term.

Hands-on project: Compute the entropy of a corpus — For a piece of text, estimate the entropy of the unigram distribution by character and by word. Then compute the cross-entropy of a small language model on the same text, and explain why perplexity = exp(cross-entropy) can compare models that use different tokenizers.

Not picked (2)
  • MacKay, Information Theory, Inference, and Learning Algorithms — A classic full-length book. Only the first few chapters are useful for ML; look up when needed.
  • Cover & Thomas, Elements of Information Theory — Aimed at communication theory; the proof density is more than this layer needs.

Optimization

Half-life: stable

optimizationbackprop-and-lossno prerequisite

Gradients, the chain rule, gradient descent and its variants (SGD, momentum, Adam), learning rate schedules, convex and non-convex. Training a network means walking down the negative gradient on a high-dimensional loss surface. This node answers: what is a gradient, how do you compute it, and how do you step without diverging. Of the four sub-topics in 数学:用到时再补, it connects most tightly to 损失与反向传播. Backpropagation is just an efficient implementation of the chain rule. The optimizer decides how those gradients become parameter updates. Learning rate warmup and cosine decay in 预训练与 scaling are also topics here.

Leads to
Loss and backpropagation
  • An overview of gradient descent optimization algorithms free

    article · 1 hour — From batch gradient descent to SGD, momentum and Adam, one piece shows what each optimizer changed from the last. You'll see what lr, momentum and betas control in training code.

Hands-on project: Write three optimizers by hand — On a 2D bowl-shaped function, hand-write the update rules for SGD, momentum and Adam, and plot the three paths. Repeat on an ill-conditioned function and explain why Adam moves more steadily.

Not picked (3)
  • Boyd & Vandenberghe, Convex Optimization — Deep learning loss surfaces are not convex, so most of this book does not apply.
  • 3Blue1Brown「Gradient descent, how neural networks learn」 — Overlaps with the video series from the same author already imported in 数学:用到时再补. After the first episode you'll keep watching anyway.
  • Goodfellow 等《Deep Learning》ch. 8 — Solid content, but it doesn't cover optimizer practice after 2016 (AdamW, warmup).

Self-check

  • Write out on paper the gradient of softmax + cross-entropy with respect to the logits (p − y), and say which chain rule each step uses.
  • Given an input of shape (batch, seq, d), write the tensor shape at each step of single-head attention, and explain why you divide by √d.
  • Explain the relationship between cross-entropy, KL divergence and perplexity in one sentence.
  • Read Adam's update formula, say what β₁, β₂ and ε each do, and say how AdamW differs from Adam.
  • Open the method section of any L2 paper and read it through without skipping a formula.

Hands-on projectDerive by hand all the gradients of a two-layer network on a single sample, then check them one by one against PyTorch autograd (attached to Math: fill in when needed).

Translated from the author's Chinese notes by a model; the Chinese page is the original.