Learning map
What to learn for the roles experienced engineers move into at AI companies, as tracks of stages and steps. The stages come in the order you need them; inside a stage, the skills come in the order this month's job posts ask for them. Each step has at most 3 links — the best free resource, a paid one only when it is clearly better, and a project or checkpoint — each with a line on why that one. Ranking never looks at commissions (how we pick).
Demand as of 2026-10-05.
Foundations
The AI groundwork every other track assumes, for engineers who already ship software: how the models work, calling them from code, search by meaning, classical ML and PyTorch, and measuring what comes out. No programming lessons; the math when a step needs it.
Stage 1How the models work
Know what a model does with your text, and where that goes wrong.
- The math, when you need it · about as needed
Pick up the linear algebra, calculus and probability a later step leans on, at the moment it leans on them.
video · 19 min — The linear algebra of a network seen rather than derived: the picture of what a layer computes that the rest of the map builds on.
book · only ch. 2–6 — Free, and written to be read in parts: linear algebra, vector calculus and probability, each chapter usable alone when a step asks for it.
Checkpoint: Backpropagation by hand, then checked — Work out the gradients of a two-layer network for one example on paper, then compare them with PyTorch's autograd.
- LLMs · 24% of technical postings (n=4,982) · about 1 week
Explain how a language model is trained and what it does, token by token, when it answers you.
video — Three and a half hours from tokens to RLHF, from one of OpenAI's founding members — the mental model every later choice rests on.
book · only ch. 2 — Chapter 2 is this skill — training data, architecture, post-training, sampling — for people who build on models; the rest is the AI Engineer track at length: evals, RAG, agents, cost.
Checkpoint: Ship one LLM feature end to end — Structured output with schema validation, streaming, a fallback model, and per-request cost and latency logged — then write down the numbers.
Stage 2Call a model from code
Get useful, checkable output from a model API and know what each call costs.
- Calling a model API · about 3 days
Make streaming, tool-using and structured calls to a model API, and account for their tokens, latency and cost.
repo — Runnable notebooks for the calls you will actually make — tool use, JSON output, prompt caching, batches — each a working example rather than a description.
video · 15 min · only 0:00–14:56 — What a token is, in a live tokenizer: why the same text costs different amounts on different models, and why models stumble on spelling and arithmetic.
Checkpoint: A client that knows what it costs — Call one model API with streaming, one tool and a JSON schema; log tokens and latency per call, and work out the cost of a thousand requests.
- Prompting · 7% of technical postings (n=4,982) · about 3 days
Improve a prompt against a test set rather than by feel.
docs — The model maker's reference, rewritten with each release so it fits the models you will call: instructions, examples, structure, output format, tools. Start with the techniques for all models.
Checkpoint: Improve one prompt with a test set, not by feel — Fifty cases with expected outputs, a baseline score, then each prompt change measured as a diff — keep the ones that move the number.
Stage 3Search by meaning
Represent text as vectors and find the passage that answers a question.
- Embeddings and vector search · about 3 days
Explain what an embedding captures, build a vector search, and say when keyword search still wins.
book — A short free book that goes from one-hot vectors to transformer embeddings with the engineering around them — the one explanation practitioners keep linking.
Checkpoint: Semantic search over your own documents — Embed a few thousand of your own documents into a vector index, then compare its top ten against keyword search on twenty queries you label.
Stage 4Machine learning, the classical way
Train and validate a model without fooling yourself, then do the same in PyTorch.
- Classical machine learning · about 1 week
Split, train, validate and pick a metric for a model on tabular data, and recognise overfitting when it happens.
course — The fundamentals — loss, generalisation, splits, classification metrics — in short interactive modules, updated by Google in 2024.
Checkpoint: A baseline you can defend — Train logistic regression and gradient-boosted trees on one tabular dataset with a proper train, validation and test split, and explain one case where accuracy is the wrong metric.
- PyTorch · 7% of technical postings (n=4,982) · about 2 weeks
Build and train a small neural network in PyTorch, from tensors and autograd up.
course — Backprop to a GPT in PyTorch, one video at a time — the fastest way for an engineer to read and debug model code.
docs — Tensors, autograd, datasets and the training loop in the framework's own words; a reference you will return to.
Checkpoint: A network from tensors up — Train a small multilayer network on a classic dataset with your own training loop, then find and fix one bug by reading its loss curve.
Stage 5Measure it
Decide whether an output is good enough before anyone else sees it.
- Evals · 15% of technical postings (n=4,982) · about 1 week
Turn error analysis on real outputs into an eval you run on every change.
article — Start from error analysis on real traces, not from a metric; then judges you check against people — the practice behind 'evaluation frameworks' in these postings.
Checkpoint: An eval harness that gates a change — Label 100 real outputs by hand, build an LLM judge and measure its agreement with you, and make CI fail when a prompt or model change drops the score.
Forward Deployed Engineer
After the foundations: build an agent on a customer's own data and tools, prove it works with their examples, run it inside their environment, and adapt the model when prompting is not enough.
Stage 1Build on the customer's data and tools
Turn a real workflow into an agent that uses the customer's systems.
- RAG / retrieval · 16% of FDE / Applied postings (n=1,078) · about 1 week
Retrieve from the customer's own documents and measure how often the right passage comes back.
article — Hybrid search (embeddings plus BM25), reranking and chunk context, each with a measured drop in retrieval failures — design with numbers.
Checkpoint: Hybrid retrieval with a labelled query set — BM25 plus vectors fused with reciprocal rank fusion, a reranker, and 50 hand-labelled queries; report recall@k and nDCG against keyword search alone.
- MCP · 3% of FDE / Applied postings (n=1,078) · about 3 days
Expose the customer's systems to the agent through MCP, with tool descriptions tuned until it picks them correctly.
docs — The protocol itself — tools, resources, transports — straight from the spec rather than from a framework's wrapper around it.
Checkpoint: An MCP server for an API you know, measured by an agent — Expose five real operations as tools, then run an agent on 30 tasks and count how often it picks the right tool with the right arguments; rewrite descriptions until that number moves.
- Agents · about 1–2 weeks
Build a tool-using agent for someone else's workflow, with step limits and error handling you can explain.
article — Workflows versus agents, and when you need neither: the handful of patterns these systems are built from, named by a lab that ships them.
course — Hands-on: tool calling, a ReAct-style loop and multi-agent setups in code you run — the gap between reading about agents and writing one.
Checkpoint: Write the agent loop yourself, then evaluate it — Messages API tool use with a step cap, tool-error retries and context trimming; score it on 30 real tasks for tool choice, argument correctness and steps taken.
Stage 2Prove it works
Show, with the customer's own examples, that it works and keeps working.
- Evals · 21% of FDE / Applied postings (n=1,078) · about 1 week
Build an eval set from the customer's real cases and report quality in their terms before and after each change.
article — Start from error analysis on real traces, not from a metric; then judges you check against people — the practice behind 'evaluation frameworks' in these postings.
Checkpoint: An eval report a customer can read — Fifty real cases from one workflow, labelled with the person who does the work today; a before-and-after table for one change, and the three failure types that remain.
Stage 3Deploy into their environment
Run it where the customer's systems and security rules are.
TypeScript · 15% of FDE / Applied postings (n=1,078) · about 3 days Skip with 5+ years as a software engineer
Read and change the TypeScript front ends and SDKs a customer integration touches.
docs — The language's own reference, short enough to read in an evening — enough to read and change the web and SDK code these roles ship to customers.
Checkpoint: Type an SDK for an API you use — Write a small typed client for a real API — request and response types, errors as values, one streaming endpoint — and publish it with its tests.
Kubernetes · 13% of FDE / Applied postings (n=1,078) · about 1 week Skip with 5+ years as a software engineer
Deploy and debug a service on a cluster you do not own.
docs — The project's own six-module tutorial: create a cluster, deploy, expose, scale and roll out an app — the vocabulary every customer's environment is written in, in an afternoon.
docs — Pods, services and running containers debugged with what you get on a cluster you do not own — describe, logs, events, a shell — the job when a customer's deployment misbehaves.
Checkpoint: Deploy your agent where a customer would run it — Package the agent as a container, deploy it to a local cluster with its secrets, a health probe and resource limits, then kill a node and watch what happens.
Stage 4Adapt the model
Know when prompting is not enough, and what to do next.
- Fine-tuning · 8% of FDE / Applied postings (n=1,078) · about 2 weeks
Fine-tune a small model for one narrow customer task and compare it with a prompted frontier model on quality, latency and cost.
course — Tokenizers, the Trainer, and fine-tuning chapters in runnable notebooks — enough to fine-tune a small open model on your own data.
docs — The library behind most supervised fine-tuning and preference tuning (DPO, GRPO) you will be asked about, with working recipes.
Checkpoint: Fine-tune against a prompted baseline — LoRA-tune a small open model on one narrow task and compare it with a prompted frontier model on the same eval set — quality, latency and cost per 1,000 calls.
AI Engineer
After the foundations: retrieval, agents and tools in a product, measured; then changing the model itself when prompting stops being enough, and serving it at a cost the product can carry.
Stage 1Build LLM features
Ship retrieval and tool use in a product, with numbers behind each.
- RAG / retrieval · 36% of AI / ML Engineer postings (n=171) · about 1 week
Retrieve the right context for a question and measure how often you do.
article — Hybrid search (embeddings plus BM25), reranking and chunk context, each with a measured drop in retrieval failures — design with numbers.
Checkpoint: Hybrid retrieval with a labelled query set — BM25 plus vectors fused with reciprocal rank fusion, a reranker, and 50 hand-labelled queries; report recall@k and nDCG against keyword search alone.
- MCP · 2% of AI / ML Engineer postings (n=171) · about 3 days
Expose a real system to agents through MCP, and measure whether they use it correctly.
docs — The protocol itself — tools, resources, transports — straight from the spec rather than from a framework's wrapper around it.
Checkpoint: An MCP server for an API you know, measured by an agent — Expose five real operations as tools, then run an agent on 30 tasks and count how often it picks the right tool with the right arguments; rewrite descriptions until that number moves.
- Agents · about 1–2 weeks
Write a tool-using agent loop yourself and measure where it fails.
article — Workflows versus agents, and when you need neither: the handful of patterns these systems are built from, named by a lab that ships them.
course — Hands-on: tool calling, a ReAct-style loop and multi-agent setups in code you run — the gap between reading about agents and writing one.
Checkpoint: Write the agent loop yourself, then evaluate it — Messages API tool use with a step cap, tool-error retries and context trimming; score it on 30 real tasks for tool choice, argument correctness and steps taken.
Stage 2Measure and improve
Know when a change made things worse, and improve past what a prompt can do.
- Evals · 51% of AI / ML Engineer postings (n=171) · about 1 week
Turn error analysis on production traces into evals that gate every change.
article — Start from error analysis on real traces, not from a metric; then judges you check against people — the practice behind 'evaluation frameworks' in these postings.
Checkpoint: An eval harness that gates a change — Label 100 real outputs by hand, build an LLM judge and measure its agreement with you, and make CI fail when a prompt or model change drops the score.
- Fine-tuning · 16% of AI / ML Engineer postings (n=171) · about 2 weeks
Fine-tune a small open model on one narrow task and prove it beats a prompted baseline.
course — Tokenizers, the Trainer, and fine-tuning chapters in runnable notebooks — enough to fine-tune a small open model on your own data.
docs — The library behind most supervised fine-tuning and preference tuning (DPO, GRPO) you will be asked about, with working recipes.
Checkpoint: Fine-tune against a prompted baseline — LoRA-tune a small open model on one narrow task and compare it with a prompted frontier model on the same eval set — quality, latency and cost per 1,000 calls.
Stage 3Train and post-train
Read, change and train model code, not only call it.
- PyTorch · 27% of AI / ML Engineer postings (n=171) · about 3 weeks
Train a small language model and change one design choice with an ablation to show for it.
course — Backprop to a GPT in PyTorch, one video at a time — the fastest way for an engineer to read and debug model code.
docs — Tensors, autograd, datasets and the training loop in the framework's own words; a reference you will return to.
Checkpoint: Train a small GPT and change one thing — Reproduce a small model with nanochat or nanoGPT on a modest dataset, change one design choice, and show the ablation with its loss curves.
- RL / post-training · 19% of AI / ML Engineer postings (n=171) · about 2 weeks
Post-train a small model with preference or verifiable rewards, and see where it games the reward.
book — Post-training as practised on language models — reward models, PPO, DPO, verifiable rewards — by a researcher who post-trains open models, free to read online.
docs — Runnable GRPO and DPO trainers, so the book's methods become an experiment on a laptop-sized model.
Checkpoint: Preference- or reward-tune a small model on a verifiable task — GRPO on arithmetic or unit-test-checked code with a small open model; plot reward against a held-out eval and explain where it starts to game the reward.
Stage 4Serve it
Run a model with the latency and cost a product needs.
- vLLM / SGLang / TensorRT · 5% of AI / ML Engineer postings (n=171) · about 1 week
Serve an open model and measure latency and throughput under load.
docs — Latency, throughput and cost of serving a model — batching, KV cache, quantization, which engine when — as one practical handbook rather than twenty blog posts.
Checkpoint: Serve a small open model and measure it — Serve a small open model with vLLM, then measure time to first token and tokens per second at 1, 8 and 32 concurrent requests, with and without quantization.