Skill · AI Hiring Index · as of 2026-10-05
RL / post-training in AI job postings
7% of technical postings at the AI companies we track mention RL / post-training (337 postings). Research Engineer postings ask for it most: 52%.
Technical postings that mention it
7%
337 postings
Which roles ask for it
Share of each technical role family's open postings that mention RL / post-training.
| Role family | Share, as a bar | Share |
|---|---|---|
| Research Engineer | 52% | |
| Research Scientist | 42% | |
| AI / ML Engineer | 19% | |
| Inference / Perf | 4% | |
| Software Eng | 3% | |
| FDE / Applied | 3% | |
| Safety / Policy | 3% | |
| Evals / Data | 3% | |
| Infra / Hardware | 1% | |
| Security | 0% |
Companies asking for it most
| Company | Postings | Of its technical postings |
|---|---|---|
| Anthropic | 40 | 12% |
| OpenAI | 36 | 7% |
| Scale AI | 28 | 27% |
| Applied Intuition | 15 | 6% |
| Thinking Machines Lab | 14 | 33% |
| Nebius | 12 | 5% |
| Fireworks AI | 9 | 29% |
| Field AI | 9 | 11% |
| Prime Intellect | 9 | 43% |
| Reflection AI | 8 | 22% |
How to learn it
- RLHF book — Nathan Lambert free
Post-training as practised on language models — reward models, PPO, DPO, verifiable rewards — by a researcher who post-trains open models, free to read online.
- TRL documentation — Hugging Face free
Runnable GRPO and DPO trainers, so the book's methods become an experiment on a laptop-sized model.
Project: Preference- or reward-tune a small model on a verifiable task — GRPO on arithmetic or unit-test-checked code with a small open model; plot reward against a held-out eval and explain where it starts to game the reward.