Layer L2
LLM core
+ A vast slice of the internet
1Pre-training
A base model that continues any text
+ Examples of good answers
2Instruction tuning
An assistant that answers
+ Preferences and rewards
3Reinforcement learning
A more helpful, careful assistant
+ Your own examples
4Fine-tuning
A model for your task
The layer
What this layer solves
Two items on the finish line: you can implement and train a GPT from scratch; you can explain the full flow from pretraining to alignment to an interviewer, with a “why” for every step. The half-life of this layer is slow. The Transformer and the pretraining objective have not changed in seven or eight years. What changes is the recipe (data, scaling, post-training methods, MoE). So review once a year. No need to chase it every quarter.
How the nodes connect into one line
The entry point is LLM: how the model works. The model does one thing: given a sequence of token, it computes the distribution of the next token. The other nodes are the parts of that one thing, ordered by data flow:
- Tokenizer: how text becomes ids. Byte-level BPE. The vocabulary sets the size of the embedding matrix and explains the model's quirks in arithmetic and spelling.
- Transformer: the model itself. attention replaces recurrence. Residuals and normalization let it stack deep. The reproduction project for this layer hangs here.
- Pretraining and scaling: training the model body into a model. The pretraining objective, data, scaling laws, and Chinchilla's “grow parameters and data together”.
- Fine-tuning: teach it your task and Post-training: shaping with rewards: post-training. SFT teaches format. RLHF / DPO shape it with preferences. Together they are “alignment”.
- Decoding and inference: picking a token from the distribution. Temperature, top-p, beam, plus why autoregression is slow and how speculative decoding helps.
- Long context and MoE and Multimodal: two kinds of change to the skeleton. The first negotiates with the quadratic cost of attention and decouples parameter count from compute. The second encodes other modalities into token and puts them in the same sequence.
Told as one line, this is the interview question: tokenizer → Transformer → pretraining (scaling) → SFT → RLHF / DPO → decoding.
What a senior engineer can cut
- Skip encoder-decoder and the BERT family. Learn decoder-only. Just know the others exist.
- Don't read the original paper first. Follow Karpathy and write it out. Then “Attention Is All You Need” becomes readable.
- For scaling laws, you only need the Chinchilla conclusion and one figure. No need to reproduce the power-law fit.
- You don't need to learn RL itself. The RLHF book covers enough of PPO / GRPO for post-training.
- Inference systems (KV cache implementation, batching, quantization) belong to L3. This layer covers concepts only.
- For multimodal, just know the two ways to wire it in. Don't chase the latest models.
Top picks
In reading order; each pick links to its node below.
- 1
Deep Dive into LLMs like ChatGPT — Andrej Karpathy free
Three and a half hours to see the whole picture, from token to RLHF. Every later node is a zoom-in on it. → LLM: how the model works
- 2
minbpe free
Write BPE yourself once. Then every tokenizer quirk has an explanation. → Tokenizer
- 3
Let's build GPT: from scratch, in code, spelled out. free
Two hours from a bigram model to multi-head attention. The first item of the finish line. → Transformer
- 4
LLMs-from-scratch free
Turns “build and train a GPT from scratch” into a checklist you can verify chapter by chapter, through instruction tuning. → Transformer
- 5
CS336: Language Modeling from Scratch free
An outline of the full pipeline. Each of the 19 lectures maps to one component, and the assignments have you implement it. → Pretraining and scaling
- 6
Training Compute-Optimal Large Language Models free
About 20 tokens per parameter. The source of every training recipe since. → Pretraining and scaling
- 7
Fine-tune a small model with a library and see what SFT looks like at the data level. → Fine-tuning: teach it your task
- 8
RLHF book — Nathan Lambert free
Reward models, PPO, DPO, verifiable rewards. The script for explaining post-training to an interviewer. → Post-training: shaping with rewards
- 9
Generation strategies — Transformers documentation free
Run greedy, sampling and beam once each. Temperature and top-p become knobs you can feel. → Decoding and inference
- 10
Mixtral of Experts free
One short paper that explains MoE routing and what “active parameters” means. → Long context and MoE
The nodes
LLM: how the model works
Half-life: slow-movingWho asks for it
Every model you can call is made in two steps. Pretraining reads a huge slice of the internet and learns one thing: predict the next token. The result is a model that can continue any text. Post-training then shows it examples of good answers and rewards the better one, and only then does it become an assistant that answers questions. Understand this and you can predict where it goes wrong: it miscounts, it forgets, it loses the thread when the context gets long, and the more confident it sounds, the more likely it is made up. If you look at only one thing on this page, make it that video: it goes from token all the way to post-training, and it skips none of the parts that explain why models get things wrong. This node is the starting point of the whole map. The other L2 nodes are its parts, and all of L4 is built on it.
- Leads to
- Decoding and inference · Embedding: turning meaning into vectors · Look up when needed: frontier reference list (≤ 15) · Weekly reads: the frontier feed (≤ 5) · Calling a model API: token, streaming, tools, cost · Serving: running models yourself · Tokenizer
- On the routes
- Foundations (How the models work)
video — Three and a half hours from tokens to RLHF, from one of OpenAI's founding members — the mental model every later choice rests on.
book · only ch. 2 — Chapter 2 is this skill — training data, architecture, post-training, sampling — for people who build on models; the rest is the AI Engineer track at length: evals, RAG, agents, cost.
Hands-on project: Ship one LLM feature end to end — Structured output with schema validation, streaming, a fallback model, and per-request cost and latency logged — then write down the numbers.
Not picked (4)
- Stanford CS336 — Already a candidate for Pretraining and scaling. It is the "hands-on version" of this node.
- Raschka「Build a Large Language Model (From Scratch)」 — Already a candidate for Transformer.
- Jay Alammar「The Illustrated GPT — 2」(2019) — Good for reading the diagrams, but Karpathy's deep-dive video covers it and is more recent.
- Hugging Face LLM Course — Already the introductory pick for Fine-tuning: teaching it your task.
Tokenizer
Half-life: slow-moving
The model never sees characters. It only sees integer ids of token. The tokenizer decides "how many token a word takes", "why Chinese costs more than English", and "why counting and reversed spelling go badly". The mainstream method is byte-level BPE: start from bytes, repeatedly merge the most common adjacent pair into a new token, and stop when the vocabulary is full. It is the first gate of LLM: how models work. In Pretraining and scaling, the vocabulary size sets the size of the embedding matrix and the output layer. In Decoding and inference, every step samples one of its ids. A mismatch between training data and tokenizer is the root of many cases of "the model suddenly got dumber".
- Builds on
- LLM: how the model works
- Leads to
- Pretraining and scaling
- minbpe free
repo · 3 hours, follow along — Minimal byte-level BPE, few hundred lines: train, encode, decode, GPT-4 regex split. Build it to explain token counts per model and why arithmetic and spelling fail. First GPT-from-scratch part.
Hands-on project: Train your own tokenizer — Use minbpe to train a 1k-vocabulary BPE on a mixed Chinese-English corpus. Compare its token count with GPT-4's cl100k on the same Chinese text, and explain where the gap comes from.
Not picked (3)
- Karpathy「Let's build the GPT Tokenizer」视频 — Already the intro pick for Calling model APIs: token, streaming, tools, cost (0:00–14:56), so not repeated here; minbpe is its code version.
- Hugging Face LLM Course ch. 6 — Covers training a tokenizer with a library; same source as the course imported by Fine-tuning: teach it your task.
- SentencePiece 论文 — Heavy on engineering detail; read it when needed, after BPE.
Transformer
Half-life: slow-moving
Since 2017 this has been the skeleton of almost every large model: hand sequence modeling to attention (each token looks at all tokens, weighted by relevance), and stack it into layers that run in parallel. It solves two problems: recurrent networks cannot train in parallel, and long-range dependencies do not carry far. The cost is that attention is quadratic in sequence length. KV cache, long context and MoE are all negotiations with this cost.
- Builds on
- Linear algebra · Loss and backpropagation · PyTorch: train a network yourself
- Leads to
- The third wave: scaling and LLM · Long context and MoE · Multimodal · Parallel training: why one card is not enough · Pretraining and scaling
video · 2 h video + 4 h coding along — From a bigram model to multi-head attention and residual blocks in two hours; each line maps to a Transformer part. After it you can write GPT from scratch, the finish line for this layer.
- nanoGPT free
repo — A 300-line minimal implementation that reproduces GPT-2 training. Reading it means understanding the pretraining loop; editing it is the cheapest ablation bench.
- LLMs-from-scratch free
repo · 2–4 h per chapter, 7 chapters — Code for Build a Large Language Model (From Scratch), a notebook per step: attention, GPT-2 weights, pretraining, finetuning. Karpathy gives intuition; this makes "build and train a GPT" a checklist.
Hands-on project: Write and train a GPT from scratch — Follow the video and hand-write a decoder-only Transformer. Train it on a small corpus until it can continue text. Then change one design choice (number of heads, position encoding, layer norm placement), run an ablation, and record the loss curve.
Not picked (4)
- Jay Alammar「The Illustrated Transformer」(2018) — The best intro for reading diagrams, but it has no code. Finish Karpathy's video first, then come back to the diagrams to consolidate.
- Vaswani 等「Attention Is All You Need」(2017) — The original paper; the entry is already in Attention Is All You Need. You can only follow it after you have written the code.
- Harvard NLP「The Annotated Transformer」 — Annotates the original paper's implementation line by line, but it is encoder-decoder, which is far from today's decoder-only mainstream.
- Stanford CS25 Transformers United — A lecture series. Good for building the full picture, not good as a first course.
Pretraining and scaling
Half-life: slow-moving
Pretraining has one objective: minimize next-token cross-entropy over as much text as possible. This node covers three things: where the data comes from, and how to deduplicate and filter it; scaling laws, meaning the power-law relation between loss and parameters, data and compute, plus Chinchilla's finding that at fixed compute the two should grow together; and the training recipe, meaning learning-rate schedule, batch size and mixed precision. It turns the Transformer and the Tokenizer into a real model. Its output goes to Fine-tuning: teach it your task and Post-training: shape it with rewards for alignment. The engineering details of compute and parallelism belong to L3. Here I only go as far as why you need that many cards.
- Builds on
- Transformer · Tokenizer · Information Theory
- Leads to
- Fine-tuning: teach it your task
course · 19×80 min, plus assignments — Only public course that builds an LLM from scratch: tokenizer, architecture, MoE, GPU, parallelism, scaling laws, data, alignment. You implement each piece. It's your outline.
paper · 2 hours — At fixed compute, scale parameters and data together, about 20 tokens per parameter. It rewrote every later training recipe. Read it to see why parameters held steady while data grew 10×.
Hands-on project: Draw a scaling curve on small models — Use nanoGPT to train three model sizes (each 4× apart in parameters) on the same corpus, each for the same number of tokens. Plot validation loss against parameter count on log axes. Then fix compute, vary the amount of data, and see whether Chinchilla's conclusion holds at toy scale.
Not picked (3)
- Kaplan 等「Scaling Laws for Neural Language Models」(2020) — The first scaling paper, but Chinchilla corrected its conclusion. Read that one first.
- Stanford CS224n — Broader (all of NLP history); pretraining gets only two lectures. CS336 fits this node better.
- Raschka「Build a Large Language Model (From Scratch)」ch. 5 — Its pretraining chapter is excellent, but the whole book is already a candidate for Transformer.
Fine-tuning: teach it your task
Half-life: slow-movingWho asks for it
Fine-tuning keeps training on your own examples so the model gets better at one narrow job: a format, a domain, a tone. Methods like LoRA train only a small add-on module instead of all the weights, so it fits on one GPU. Use it once prompting and retrieval have stopped improving things, and you need an eval that proves it beats the prompt-only model. Fine-tuning without a baseline comparison doesn't count.
- Builds on
- PyTorch: train a network yourself · Evals: knowing whether it got better · Pretraining and scaling
- Leads to
- Post-training: shaping with rewards
- On the routes
- Forward Deployed Engineer (Adapt the model) · AI Engineer (Measure and improve)
course — Tokenizers, the Trainer, and fine-tuning chapters in runnable notebooks — enough to fine-tune a small open model on your own data.
docs — The library behind most supervised fine-tuning and preference tuning (DPO, GRPO) you will be asked about, with working recipes.
Hands-on project: Fine-tune against a prompted baseline — LoRA-tune a small open model on one narrow task and compare it with a prompted frontier model on the same eval set — quality, latency and cost per 1,000 calls.
Not picked (3)
- Ouyang 等「Training language models to follow instructions with human feedback」(InstructGPT, 2022) — The original source for the three-stage SFT → RM → PPO pipeline. The concept is already covered by the RLHF book in Post-training: shaping with rewards.
- Hu 等「LoRA」(2021) — The original paper is short, but the TRL / PEFT docs already turn it into a usable recipe.
- Unsloth 文档 — A tooling item that changes fast. Not picked as a principles pick.
Post-training: shaping with rewards
Half-life: slow-movingWho asks for it
Reinforcement learning learns from scores, not from examples. The model tries a few answers, a reward says which is better (human preference, or a check like "passes the tests"), and training pushes it toward the higher-scoring ones. This is how a base model becomes a useful assistant. It is also how reasoning models learn to solve problems step by step. If a reward can be gamed, the model will learn to game it. Finding that out is part of the work. Understand this and you know where "alignment," "refusals," and "sycophancy" come from.
- Builds on
- Fine-tuning: teach it your task
- On the routes
- AI Engineer (Train and post-train)
book — Post-training as practised on language models — reward models, PPO, DPO, verifiable rewards — by a researcher who post-trains open models, free to read online.
docs — Runnable GRPO and DPO trainers, so the book's methods become an experiment on a laptop-sized model.
Hands-on project: Preference- or reward-tune a small model on a verifiable task — GRPO on arithmetic or unit-test-checked code with a small open model; plot reward against a held-out eval and explain where it starts to game the reward.
Not picked (4)
- Rafailov 等「Direct Preference Optimization」(2023) — The original DPO paper. The RLHF book has a full chapter on it, so read the book first, then the paper.
- Lilian Weng「Reinforcement Learning from Human Feedback」/ 早期 RL 系列 — High-quality survey, but the RLHF book is newer and more complete.
- Sutton & Barto, Reinforcement Learning: An Introduction — A classic RL textbook. LLM post-training uses only a small corner of it. Not worth the whole book.
- Stanford CS336 对齐三讲 — Already in this layer through a candidate from Pretraining and scaling.
Decoding and inference
Half-life: slow-moving
A trained model only gives a distribution over the next token. Choosing a token from that distribution is decoding: greedy, beam search, sampling with temperature, top-k / top-p truncation. Temperature changes how flat the distribution is. Top-p cuts off the long tail. Beam looks for the most likely whole sequence, not the most likely step. On the inference side the cost is autoregression: every token generated needs a pass through the whole model, and the bottleneck is memory bandwidth. KV cache and speculative decoding are two answers to that. This node covers the concepts. Serving systems (batching, quantization, scheduling) are in L3, Serving: running models yourself. It uses the vocabulary of Probability and statistics directly and builds on LLM: how the model works.
- Builds on
- LLM: how the model works · Probability and Statistics
- Leads to
- Serving: running models yourself
docs · 1 hour — Greedy, sampling, beam: what each is, which tasks fit, a runnable generate() call each; ends with a link to a top-k / top-p article. Run them: temperature and top_p become knobs, not nouns.
paper · 1 hour — Small model drafts a few tokens, big model checks them in one pass: same output distribution, 2–3x faster. Shows decoding is memory-bound, not compute-bound; a top interview question.
Hands-on project: Write the sampling loop by hand — Take a small open-source model. Without generate(), write your own token-by-token loop that implements temperature, top-k and top-p. Then implement the simplest speculative decoding (small model drafts 4 tokens, big model verifies) and measure tokens per second for all three.
Not picked (3)
- Hugging Face「How to generate text」博客 (2020) — The well-known article on top-k / top-p, but the docs page already links it, and HN praise isn't enough.
- Holtzman 等「The Curious Case of Neural Text Degeneration」(2019) — The original top-p paper; look up when needed.
- Lilian Weng「Large Transformer Model Inference Optimization」 — The content belongs to L3.
Long context and MoE
Half-life: slow-moving
Both directions are negotiations with the cost of the Transformer. Long context: attention is quadratic in sequence length, and position encodings break down beyond the training length. That is why there are rotary position embeddings (RoPE) and their extrapolation, sparse / sliding-window attention, and the lost-in-the-middle effect, where a longer context does not mean the model uses it well. MoE: replace each layer's FFN with several experts and let a router activate only a few of them. This decouples parameter count from per-token compute, so a large model can hold more knowledge at the same inference cost. Both build on the Transformer. The training-side parallelism and the inference-side memory cost belong to L3 (Serving: running models yourself).
- Builds on
- Transformer
- Mixtral of Experts free
paper · 1 hour — Shortest paper that makes MoE clear: 8 experts per layer, 2 picked per token, 47B params but 13B compute. Explains why total and active parameters differ, and what MoE buys and costs (memory).
Hands-on project: Add an MoE layer to nanoGPT — Replace the FFN in the [[transformer]] project with 4 experts and top-2 routing, and add a load-balancing loss. Compare validation loss and memory use at the same compute. Then raise the context from 256 to 2048 and record how attention memory and speed change.
Not picked (4)
- Hugging Face「Mixture of Experts Explained」(2023) — Friendlier explanation, but its HN score of 29 misses the bar. Keep it as a companion to the Mixtral paper.
- Fedus 等「Switch Transformers」(2021) — The founding MoE paper, and long. The Mixtral paper alone is enough to build the model.
- Su 等「RoFormer」(2021) — The original RoPE paper. Read it when you change position encoding in the Transformer project.
- Liu 等「Lost in the Middle」(2023) — The phenomenon matters, but the conclusion fits in one sentence.
Multimodal
Half-life: slow-moving
There are only two mainstream ways to let a language model see images, hear audio and watch video. One is to encode the other modality into tokens and put them in the same sequence (the LLaVA line). The other is to connect them into the middle layers of the language model with cross-attention. Both need a pretrained modality encoder (for vision, usually a CLIP-style contrastive model) and a stage of "alignment" training. Learn this, and in an interview you can answer "how does a multimodal model differ from a text-only model" in terms of structure, not products. It is built on the Transformer. The convolution / ViT basics of the vision encoder are in Classic architectures: CNN, RNN / LSTM, residuals. The instruction tuning from Fine-tuning: teaching it your task shows up here once more, in the form of image-text pairs.
paper · 1.5 hours — Original simplest multimodal LLM recipe: a projection layer turns patch features into image tokens for the LM, then instruction tuning. You can then explain how a model sees in three sentences.
Hands-on project: Attach a vision encoder to a small language model — Use an off-the-shelf CLIP vision encoder and an open-source language model under 1B parameters. Train only a linear projection layer on a small image-text dataset for image captioning. Record what it gets right and wrong, and explain what the projection layer is aligning.
Not picked (3)
- Raschka「Understanding Multimodal LLMs」(2024) — The clearest survey comparing the two architectures, but its independent endorsement falls short (HN: 4 points). Strongly recommended as a companion after LLaVA.
- Radford 等 CLIP 论文 (2021) — The source of the vision encoder and heavily cited, but it is contrastive learning, not a multimodal LLM. Look up when needed.
- Chip Huyen「Multimodality and Large Multimodal Models」(2023) — Endorsement is also too thin, and the content overlaps with Raschka.
Self-check
- Without a reference, write the forward pass of a decoder-only Transformer (embedding, position, multi-head causal attention, FFN, residuals and normalization). Train it on a small corpus until it can continue text.
- Explain why attention is divided by √d, why it needs a causal mask, and what changes if layer norm goes before or after.
- Describe the full flow of pretraining → SFT → RLHF / DPO out loud. For each step, say what it changes in the model and what data it uses.
- State the Chinchilla conclusion, and why it rewrote the parameter-to-data ratio of later models.
- Explain what temperature, top-p and speculative decoding each do, and why MoE “total parameters” and “active parameters” are two different numbers.
Hands-on projectWrite and train a GPT from scratch, then change one design choice as an ablation and record the loss curve (attached to Transformer).
Translated from the author's Chinese notes by a model; the Chinese page is the original.