Skip to main content

Role · AI Hiring Index · as of 2026-10-05

AI Engineer jobs at AI companies

Sierra (27) and Nebius (17) lead 171 open AI Engineer roles across 42 AI companies. 9% are remote and 54% are in the Bay Area.

AI engineers build the product features that run on language models and the systems underneath them — retrieval, agents, evaluation, and fine-tuning where prompting is not enough. At model companies many also train and serve models, which is why PyTorch appears in so many of these postings.

Product engineering at AI-native companies; applied research or platform teams at the labs.

Demand this month

Open roles
171
2% of all openings
New in the last 7 days
171
Remote
9%
Bay Area
54%

What AI Engineer postings ask for

Share of 171 open postings that mention each skill.

SkillShare, as a barShare
LLMs73%n=124
Python54%n=93
Evals51%n=88
RAG / retrieval36%n=61
TypeScript28%n=48
PyTorch27%n=46
Prompting26%n=45
Go25%n=43
RL / post-training19%n=32
Fine-tuning16%n=28
C++12%n=20
Kubernetes11%n=18
Distributed training9%n=15
CUDA / Triton6%n=11
JAX6%n=10
vLLM / SGLang / TensorRT5%n=8
Rust4%n=6
MCP2%n=3
PhD2%n=3

Posted base pay

Median posted range $194K–$324K base, from 102 postings that publish pay (60% of this role's openings). This is the base range a posting advertises — not total compensation, which adds equity and bonus; for that, Levels.fyi is the reference.

Midpoint of each posted base range, in thousands of US dollars: <150: 0; 150–200: 11; 200–250: 33; 250–300: 20; 300–350: 32; 350–400: 2; 400–450: 1; 450+: 3

Half of the posted ranges have their midpoint between $221K and $300K; the median midpoint is $265K.

What to learn

What you can skip

Backend and API design, data modelling, testing, deployment and on-call — the software engineering these roles are built on. What changes is that model behaviour is now part of the system you test.

AI Engineer

After the foundations: retrieval, agents and tools in a product, measured; then changing the model itself when prompting stops being enough, and serving it at a cost the product can carry.

  1. Stage 1Build LLM features

    Ship retrieval and tool use in a product, with numbers behind each.

    1. RAG / retrieval · 36% of AI / ML Engineer postings (n=171) · about 1 week

      Retrieve the right context for a question and measure how often you do.

      • Introducing Contextual Retrieval — Anthropic free

        article — Hybrid search (embeddings plus BM25), reranking and chunk context, each with a measured drop in retrieval failures — design with numbers.

      Checkpoint: Hybrid retrieval with a labelled query set — BM25 plus vectors fused with reciprocal rank fusion, a reranker, and 50 hand-labelled queries; report recall@k and nDCG against keyword search alone.

    2. MCP · 2% of AI / ML Engineer postings (n=171) · about 3 days

      Expose a real system to agents through MCP, and measure whether they use it correctly.

      • Model Context Protocol — introduction and spec free

        docs — The protocol itself — tools, resources, transports — straight from the spec rather than from a framework's wrapper around it.

      Checkpoint: An MCP server for an API you know, measured by an agent — Expose five real operations as tools, then run an agent on 30 tasks and count how often it picks the right tool with the right arguments; rewrite descriptions until that number moves.

    3. Agents · about 1–2 weeks

      Write a tool-using agent loop yourself and measure where it fails.

      • Building effective agents — Anthropic free

        article — Workflows versus agents, and when you need neither: the handful of patterns these systems are built from, named by a lab that ships them.

      • Hugging Face Agents Course free

        course — Hands-on: tool calling, a ReAct-style loop and multi-agent setups in code you run — the gap between reading about agents and writing one.

      Checkpoint: Write the agent loop yourself, then evaluate it — Messages API tool use with a step cap, tool-error retries and context trimming; score it on 30 real tasks for tool choice, argument correctness and steps taken.

  2. Stage 2Measure and improve

    Know when a change made things worse, and improve past what a prompt can do.

    1. Evals · 51% of AI / ML Engineer postings (n=171) · about 1 week

      Turn error analysis on production traces into evals that gate every change.

      Checkpoint: An eval harness that gates a change — Label 100 real outputs by hand, build an LLM judge and measure its agreement with you, and make CI fail when a prompt or model change drops the score.

    2. Fine-tuning · 16% of AI / ML Engineer postings (n=171) · about 2 weeks

      Fine-tune a small open model on one narrow task and prove it beats a prompted baseline.

      • Hugging Face LLM Course free

        course — Tokenizers, the Trainer, and fine-tuning chapters in runnable notebooks — enough to fine-tune a small open model on your own data.

      • TRL documentation — Hugging Face free

        docs — The library behind most supervised fine-tuning and preference tuning (DPO, GRPO) you will be asked about, with working recipes.

      Checkpoint: Fine-tune against a prompted baseline — LoRA-tune a small open model on one narrow task and compare it with a prompted frontier model on the same eval set — quality, latency and cost per 1,000 calls.

  3. Stage 3Train and post-train

    Read, change and train model code, not only call it.

    1. PyTorch · 27% of AI / ML Engineer postings (n=171) · about 3 weeks

      Train a small language model and change one design choice with an ablation to show for it.

      • Neural Networks: Zero to Hero — Andrej Karpathy free

        course — Backprop to a GPT in PyTorch, one video at a time — the fastest way for an engineer to read and debug model code.

      • Learn the Basics — PyTorch tutorials free

        docs — Tensors, autograd, datasets and the training loop in the framework's own words; a reference you will return to.

      Checkpoint: Train a small GPT and change one thing — Reproduce a small model with nanochat or nanoGPT on a modest dataset, change one design choice, and show the ablation with its loss curves.

    2. RL / post-training · 19% of AI / ML Engineer postings (n=171) · about 2 weeks

      Post-train a small model with preference or verifiable rewards, and see where it games the reward.

      • RLHF book — Nathan Lambert free

        book — Post-training as practised on language models — reward models, PPO, DPO, verifiable rewards — by a researcher who post-trains open models, free to read online.

      • TRL documentation — Hugging Face free

        docs — Runnable GRPO and DPO trainers, so the book's methods become an experiment on a laptop-sized model.

      Checkpoint: Preference- or reward-tune a small model on a verifiable task — GRPO on arithmetic or unit-test-checked code with a small open model; plot reward against a held-out eval and explain where it starts to game the reward.

  4. Stage 4Serve it

    Run a model with the latency and cost a product needs.

    1. vLLM / SGLang / TensorRT · 5% of AI / ML Engineer postings (n=171) · about 1 week

      Serve an open model and measure latency and throughput under load.

      • LLM Inference Handbook — Modular free

        docs — Latency, throughput and cost of serving a model — batching, KV cache, quantization, which engine when — as one practical handbook rather than twenty blog posts.

      Checkpoint: Serve a small open model and measure it — Serve a small open model with vLLM, then measure time to first token and tokens per second at 1, 8 and 32 concurrent requests, with and without quantization.

One project

An LLM feature with its own eval harness — Build one feature on retrieval or tool use, label a test set by hand, gate changes on it in CI, and publish the error analysis — the portfolio piece hiring managers ask you to walk through.

The AI Engineer track on the learning map

Open AI Engineer roles

See all 174 postings

Posted demand, not hires: open postings on the public job boards of the AI companies we track, read every night. A repost counts once; a role family is assigned from the title; skill tags are checked against the postings every month and a skill below 90% precision is not shown; pay is the posted base range, shown only when at least 10 postings publish it. How we count.