Skip to main content

Layer L3

Systems and Infrastructure

The layer

Where this layer ends for application engineers

L3 covers what the model runs on, how to make it fast, and what one run costs. For an application engineer working mostly at L4, the finish line for this layer is "can explain it + deployed once". You can explain the trade-offs of an inference service in a system design interview. You can run an open-source model yourself and measure the numbers. You do not need to write CUDA, and you do not need to have run distributed training. Inside the boundary there are four topics: GPU basics: why it's the one, and where the bottleneck is (why the hardware is the way it is), Serving: running models yourself (inference serving, the center of this layer), Parallel training: why one card is not enough (why training needs thousands of cards), and Compute economics: where training money goes, how inference is billed (how the money works). Add one more layer for landing it: Kubernetes: running it where the customer is, the vocabulary for finally deploying the model and the application into a customer's environment.

How the four quantities squeeze each other

Inference serving has only four quantities: throughput, latency, memory and cost. They are not four independent metrics. They are several ends of one lever. Every decode step reads all the weights from memory once, so a single request barely uses the card. Merge requests into a batch and the weights are read once to feed many requests. Throughput goes up, but the gap between tokens for each request gets longer, so latency gets worse. How big a batch you can form depends on how much memory the KV cache has left, and the KV cache grows with both concurrency and context length. Cost is the inverse of throughput: the card's rental price is fixed, so how many tokens you get out of one hour decides the price per million tokens. So every optimization falls into three things. Quantization makes the weights smaller (a bandwidth bottleneck gets faster directly, and the same memory holds more concurrency). Paging and prefix caching waste less KV. Parallelism lets a model that does not fit on one card fit. If you can explain this section, you can answer every system design question in this layer.

Two lines

Inference side (the center of this layer): start from the "compute-bound or bandwidth-bound" check in GPU basics: why it's the one, and where the bottleneck is. Go on to continuous batching, paged KV, quantization and parallelism in Serving: running models yourself. The end point is to deploy once yourself and measure the four quantities, then turn the cost per million tokens into money in Compute economics: where training money goes, how inference is billed.

Training side (you only need to be able to explain it): one GPU cannot hold a big model's parameters, gradients and optimizer state. So Parallel training: why one card is not enough covers four ways to split it: data parallelism copies the model, ZeRO spreads the state, tensor parallelism splits a single layer, and pipeline parallelism splits segments of layers. It also covers where each split adds communication. This explains why training takes thousands of cards. It also explains where the training cost column in Compute economics: where training money goes, how inference is billed comes from.

Top picks

In reading order; each pick links to its node below.

  1. 1

    What Every Developer Should Know About GPU Computing free

    Know what a GPU looks like first, then you can explain why it swallows matrix multiplies so well → GPU basics: why it's used, and where the bottleneck is

  2. 2

    Making Deep Learning Go Brrrr From First Principles free

    How to tell compute, bandwidth and overhead bottlenecks apart. Every later trade-off depends on it → GPU basics: why it's used, and where the bottleneck is

  3. 3

    How is LLaMa.cpp possible? free

    Applies the bottleneck check to LLM decode and works it out. Shows why quantization buys speed directly → GPU basics: why it's used, and where the bottleneck is

  4. 4

    LLM Inference Handbook — Modular free

    A full handbook on inference serving, with one section for each of the four quantities and each lever → Serving: running models yourself

  5. 5

    Inside vLLM: Anatomy of a High-Throughput LLM Inference System free

    Follow one request through a real engine. This is the diagram you draw in a system design interview → Serving: running models yourself

  6. 6

    vLLM documentation free

    The engine and measurement tools for the deployment project. This is where the numbers come from → Serving: running models yourself

  7. 7

    The Ultra-Scale Playbook: Training LLMs on GPU Clusters free

    Introduces the four kinds of parallelism one by one: what each saves and what extra communication each adds → Parallel training: why one card is not enough

  8. 8

    How to Scale Your Model: A Systems View of LLMs on TPUs free

    Uses the same roofline ruler to work out how many chips training and inference need → Parallel training: why one card is not enough

  9. 9

    DeepSeek-V3/R1 Inference System Overview free

    An inference ledger with real numbers. A first-hand example of how cost per token is calculated → The economics of compute: where training money goes, how inference is priced

  10. 10

    Learn Kubernetes Basics — Kubernetes documentation free

    The vocabulary you need when the model finally gets deployed into a customer's environment → Kubernetes: running it at the customer's site

The nodes

GPU basics: why it's used, and where the bottleneck is

Half-life: fast-moving

ml-mathgpu-basicscompute-economicsparallel-trainingserving

Deep learning is mostly matrix multiplication. A GPU uses thousands of simple cores to do the same arithmetic in parallel, which fits it exactly. You need to be able to explain three quantities: compute (FLOPs/s, set by core count and tensor cores), memory capacity (how large a set of weights and KV cache fits), and memory bandwidth (how many bytes per second can move into the compute units). There are only two estimation rules. One forward pass costs about two floating-point operations per parameter, so a 7B model needs about 14 GFLOPs per token. In decode, every generated token reads all the weights once, so the speed ceiling ≈ bandwidth ÷ weight bytes. Compare these two numbers with the chip's specs and you know whether you are compute-bound or bandwidth-bound, and which side quantization and batching each help. The boundary for this layer: be able to calculate and judge, not write CUDA.

Builds on
Math: learn it when you need it
Leads to
The economics of compute: where training money goes, how inference is priced · Parallel training: why one card is not enough · Serving: running models yourself
  • What Every Developer Should Know About GPU Computing free

    article · 40 minutes — GPU hardware intro (SM, warp, register and memory hierarchy) at application-engineer depth. Explains why GPUs handle matrix multiplies and why data movement is the bottleneck.

  • Making Deep Learning Go Brrrr From First Principles free

    article · 45 minutes — How to tell the three bottlenecks apart (compute, bandwidth, overhead), with arithmetic intensity and FLOPs estimates. If asked in an interview 'where is this model slow', this is the framework.

  • How is LLaMa.cpp possible? free

    article · 30 minutes — Applies the first two to LLM decode: weight bytes ÷ bandwidth = per-token floor. You can work out why quantization buys speed and why a laptop runs 7B.

Not picked (4)
  • GPU MODE(原 CUDA MODE)讲座系列 — Written for people who write kernels; deeper than this layer's 'can explain, needn't write' line. GitHub 6,819 stars (as of 2026-10-08), kept as an entry point if you want to dig down.
  • Transformer Inference Arithmetic(kipply,2022) — Overlaps with the finbarr piece; reached under 100 points on HN both times it made the front page.
  • How to Scale Your Model 第 1 章(rooflines)与第 12 章(GPUs) — The same book is already a pick in Parallel training: why one card is not enough; no duplicate slot.
  • AI Engineering 第 9 章(Chip Huyen) — Strong on inference optimization, but the same book is a paid pick at LLM: how models work, and the handbook in Serving: run your own model covers it.

Serving: running models yourself

Half-life: fast-movingWho asks for it

requestsbatchGPUtokens outBigger batches: more tokens per second in all,and a longer wait for each first token

A hosted API runs the model for you. Run an open-source model yourself and the job becomes GPU time: batch many requests together, cache the repeated parts, and trade off latency (how long until the first token) against throughput (tokens per second across all users). Inference engines like vLLM and SGLang do most of the work. The skill in serving is knowing which of their settings to change, and measuring under a real load. The goal is not to write CUDA. It is to explain, in a system design question, how throughput, latency, GPU memory and cost squeeze each other.

What determines the four quantities

  • Throughput (tokens produced per second) mostly comes from batching. In the decode phase every step reads all the weights once. Serve one request at a time and memory bandwidth is mostly wasted. Merge dozens of requests into one batch and one weight read feeds many of them, so throughput grows roughly linearly until compute or KV memory hits its ceiling first. Continuous batching lets new requests join a running batch at any time, without waiting for the whole batch to finish.
  • Latency has two parts. Time to first token is set by prefill (computing the whole prompt in one pass) and scales with prompt length and queue depth. After that, the gap between tokens is set by the time of one decode step. A bigger batch gives higher throughput but also stretches each request's gap between tokens. Throughput and latency are two ends of the same lever.
  • GPU memory holds two things: weights and the KV cache. Weights are a fixed amount. The KV cache grows with concurrency × context length. When it doesn't fit you have to shrink the batch, and throughput drops with it. Paged management (PagedAttention) cuts waste, and prefix caching keeps a shared system prompt from taking space twice.
  • Cost is the inverse of throughput. The hourly rent of a card is fixed, so the price per million tokens depends only on how many tokens that hour produced. So the levers for cutting cost are the first three: a bigger batch, less KV memory, quantization (smaller weights, so a bandwidth-bound step gets faster and the same memory fits more concurrency), and parallelism (a model too big for one card fits, or compute stacks up). Tensor parallelism splits a single layer within a node and lowers latency. Pipeline parallelism spans nodes and buys throughput. Both add communication.

Spec for the hands-on deployment

The spec only. Do not execute it here. The numbers from the run go into the table below.

What to deploy. One candidate per tier, from the same family so they can be compared: Qwen3-8B for the 7B class (Hugging Face has official GGUF, AWQ and community MLX 4-bit versions, so none of the three paths needs your own conversion), and Qwen3-1.7B for the small model. If a newer 7–9B dense model from the same family exists when you run it, swap it in. The spec stays the same.

Where to run it, two paths:

Path Tool Fits
Local Apple Silicon llama.cpp (llama-bench measures throughput, llama-server provides the API) or MLX (mlx_lm.generate prints prompt / generation tokens/s and peak memory directly) Zero cost, easy to retry. The effect of quantization on speed and memory is most visible here. Concurrency tests stop at 8. On unified memory there is no "out of GPU memory", only "slower"
Rent a cloud GPU for one hour One 24–80 GB card, start the service with vllm serve, load it with vllm bench serve See the real batching curve, the memory use of paged KV, and the rental cost per million tokens. Concurrency 1 / 8 / 32 all fit within the hour

What to measure, recorded for every configuration:

  1. Throughput: output tokens/s (whole machine), and tokens/s for a single request.
  2. Time to first token (TTFT) p50 / p90; time per output token (TPOT) p50 / p90. The vLLM bench gives the median and P99 by default. Compute p90 from the raw results or set it with the percentile parameter. On the local path, time it yourself with the same load script.
  3. Memory use: the weights part + the KV cache part (the vLLM startup log gives the KV reservation, nvidia-smi gives the total, and locally you read peak memory).
  4. Cost per million tokens: cloud GPU = hourly rent ÷ (output tokens/s × 3600) × 10⁶. Locally, record 0 and note that electricity is not counted.

Variable matrix: model (8B / 1.7B) × precision (bf16 or fp16 / 4-bit quantized) × concurrency (1 / 8 / 32) × path (local / cloud), with a fixed prompt length (for example 512 in, 128 out). On the local path, skip concurrency 32.

Output. A one-page comparison table: one row per configuration, with the four groups of numbers above as columns. Under the table, three sentences of conclusions: how much speed quantization bought, how many times throughput rose from concurrency 1 to 32 and how much TPOT got worse, and what a million tokens costs in the cloud, set next to the list price of one hosted API.

Builds on
LLM: how the model works · GPU basics: why it's used, and where the bottleneck is · Decoding and inference
Leads to
The economics of compute: where training money goes, how inference is priced · Kubernetes: running it at the customer's site
On the routes
AI Engineer (Serve it)
  • LLM Inference Handbook — Modular free

    docs — Latency, throughput and cost of serving a model — batching, KV cache, quantization, which engine when — as one practical handbook rather than twenty blog posts.

  • Inside vLLM: Anatomy of a High-Throughput LLM Inference System free

    article · 90 min — One request followed through the scheduler, paged KV cache, continuous batching, prefix caching and speculative decoding, ending on the latency-against-throughput curve a design answer needs.

  • vLLM documentation free

    docs — The engine the cloud-GPU route runs on: the docs explain the trade-off behind each setting, and the bundled vllm bench serve prints TTFT, TPOT and throughput — the deployment project's numbers.

Hands-on project: Serve a small open model and measure it — Serve a small open model with vLLM, then measure time to first token and tokens per second at 1, 8 and 32 concurrent requests, with and without quantization.

Not picked (6)
  • SGLang 文档 — The second mainstream engine. Its concepts mirror vLLM's (RadixAttention is prefix caching), so one per deployment is enough. GitHub 36,861 stars (taken 2026-10-08).
  • TensorRT — LLM docs — worth it only when chasing peak performance on the NVIDIA stack. The deployment bar is higher than this layer.
  • llama.cpp、MLX — Tools you use directly on a local machine. They appear in the deployment spec and take no pick slot. Both had commits that day.
  • Life of an inference request (vLLM V1)(Ubicloud,2025 — 06, HN 175 points) — covers the same path as Gordić's piece but shorter. Pick the more complete of the two.
  • Nano — vLLM (2025–2026) — a good way to learn engines by reading source, but beyond the boundary of "can explain it + deployed it once".
  • AI Engineering 第 9 章(Chip Huyen) — The same book is already a paid pick in the LLM: how the model works node. Its inference optimization content overlaps with the handbook.

Parallel training: why one card is not enough

Half-life: fast-moving

gpu-basicstransformerparallel-training

The parameters, gradients and optimizer states of a 70B model already take terabytes in mixed precision, and one card has less than 200 GB of memory. So training starts as a question of how to cut the model up and where to put the pieces. Data parallelism keeps a full copy of the model on each card, lets each compute its own batch, then syncs gradients. ZeRO slices optimizer states, gradients and parameters in turn, so memory is no longer copied per card. Tensor parallelism splits one layer's matrices across several cards and needs communication in every layer, so I only use it inside a node. Pipeline parallelism cuts the model into stages by layer and passes activations along, and the cost is bubbles. Thousands of cards come from multiplying several kinds of parallelism together. The hard engineering part is hiding communication behind computation. This layer only asks that you can explain what each kind of parallelism splits and where the communication happens. It does not ask you to have run any of it.

Builds on
GPU basics: why it's used, and where the bottleneck is · Transformer
  • The Ultra-Scale Playbook: Training LLMs on GPU Clusters free

    book · One weekend — From single-card memory to data, ZeRO, tensor, pipeline, context and expert parallelism: what each saves, what communication it adds. You can explain why training needs thousands of cards.

  • How to Scale Your Model: A Systems View of LLMs on TPUs free

    book · One weekend — One roofline yardstick for training and inference. Communication and compute per strategy are formulas, so you can derive chip count and best split on a whiteboard. Ch. 12 moves it to GPU.

  • CS336: Language Modeling from Scratch free

    course · 4 systems lectures, ~6 hours — For listening, not reading: lectures 5 (GPU), 7–8 (parallelism), 10 (inference), by researchers on design trade-offs; covers this layer's four nodes. Videos are in the YouTube playlist.

Not picked (3)
  • picotron / nanotron(Hugging Face) — Companion code for the playbook. It runs, but this layer stops at "can explain" and does not require hands-on training.
  • Megatron — Official docs for LM and DeepSpeed: implementation docs, useful once you already use them; the playbook explains the concepts more clearly.
  • GPU MODE 的 NCCL / 集合通信讲座 — Too deep for this layer; same series listed under Not picked in GPU basics: why it is the one, and where the bottleneck is.

The economics of compute: where training money goes, how inference is priced

Half-life: fast-moving

gpu-basicsservingcompute-economics

Training cost is one-off: number of cards × hours × unit price, plus failed reruns, data and people. Inference cost is ongoing, and it decides the product's gross margin. Cost per million token has one formula: GPU rent per hour ÷ the number of token actually produced in that hour. So throughput is the inverse of cost, and throughput is set by batching size, KV cache hit rate, quantization and parallelism. This is exactly the last of the four quantities in Serving: running models yourself. On the supply side, look at three things: chips (how much of the price is compute and how much is HBM, and who holds the capacity), power and data centers, and the hourly price in the rental market. The aim of this section is that, in a system design question, you can work out "how much will this feature burn in a month" and name the four levers for cutting cost.

Builds on
GPU basics: why it's used, and where the bottleneck is · Serving: running models yourself
  • DeepSeek-V3/R1 Inference System Overview free

    article · 30 minutes — A model lab's own inference ledger: nodes, daily tokens, daily GPU cost, revenue at list price. The only first-hand per-token cost with real numbers; why big batching and expert parallelism cut cost.

  • Memory has grown to nearly two-thirds of AI chip component costs free

    article · 15 minutes — Key supply-side fact: HBM, not logic, is now the biggest cost in an AI chip. Explains why memory bandwidth is pricey and why compute growth is stuck on memory supply, not just wafers.

Not picked (4)
  • How AI Labs Are Solving the Power Crisis(SemiAnalysis,2025 — 12) — The most detailed treatment of the power supply line, but the cost analysis is behind a paywall, HN 167 points; a paid extra if needed.
  • Inference Economics of Language Models(Epoch AI,2025 — 06) — A good piece on modeling inference cost, but no independent endorsement (HN 4 points); kept as further reading.
  • LLM Inference Economics from First Principles(Tensor Economics,2025 — 05) — Same: careful derivation, no endorsement.
  • AI's $600B Question(Sequoia,2024) — It is about return on investment, not unit economics, so it is far from the finish line for this layer.

Kubernetes: running it at the customer's site

Half-life: fast-movingWho asks for it

FDE work often lands in the customer's cloud, on a cluster you don't own. If you have shipped services before, you already know most of this: deployment, config and secrets, resource limits, rolling updates, reading logs and metrics. You don't need to know how to operate the cluster itself. Everything you did on k3s counts.

Builds on
Serving: running models yourself
On the routes
Forward Deployed Engineer (Deploy into their environment)
  • Learn Kubernetes Basics — Kubernetes documentation free

    docs — The project's own six-module tutorial: create a cluster, deploy, expose, scale and roll out an app — the vocabulary every customer's environment is written in, in an afternoon.

  • Troubleshooting Applications — Kubernetes documentation free

    docs — Pods, services and running containers debugged with what you get on a cluster you do not own — describe, logs, events, a shell — the job when a customer's deployment misbehaves.

Hands-on project: Deploy your agent where a customer would run it — Package the agent as a container, deploy it to a local cluster with its secrets, a health probe and resource limits, then kill a node and watch what happens.

Self-check

  • Given a card's compute and bandwidth, and a model's parameter count and precision, I can work out its decode tokens/s ceiling in my head, and say whether it is compute-bound or bandwidth-bound right now.
  • I can explain how throughput, TTFT, TPOT and KV memory each change as concurrency goes from 1 to 32, and which of them quantization touches.
  • Given a product's daily token volume and one card's rental price, I can work out the cost per million tokens and the monthly bill, and name the four levers for cutting cost.
  • I can explain what data parallelism, ZeRO, tensor parallelism and pipeline parallelism each split, where the communication happens, and why training needs thousands of cards.
  • I have deployed an open-source model myself and produced a one-page comparison table with numbers.

Hands-on projectRun Qwen3-8B and Qwen3-1.7B on a local machine (llama.cpp / MLX) or on a cloud GPU for one hour (vLLM). At 1 / 8 / 32 concurrency, with and without quantization, measure throughput, p50 / p90 of TTFT and TPOT, memory use, and cost per million tokens. Produce a one-page comparison table. The spec is in Serving: running models yourself, under "the hands-on deployment spec".

Translated from the author's Chinese notes by a model; the Chinese page is the original.