Skip to main content

Skill · AI Hiring Index · as of 2026-10-05

RL / post-training in AI job postings

7% of technical postings at the AI companies we track mention RL / post-training (337 postings). Research Engineer postings ask for it most: 52%.

Technical postings that mention it
7%
337 postings

Which roles ask for it

Share of each technical role family's open postings that mention RL / post-training.

Role familyShare, as a barShare
Research Engineer52%
Research Scientist42%
AI / ML Engineer19%
Inference / Perf4%
Software Eng3%
FDE / Applied3%
Safety / Policy3%
Evals / Data3%
Infra / Hardware1%
Security0%

Companies asking for it most

CompanyPostingsOf its technical postings
Anthropic4012%
OpenAI367%
Scale AI2827%
Applied Intuition156%
Thinking Machines Lab1433%
Nebius125%
Fireworks AI929%
Field AI911%
Prime Intellect943%
Reflection AI822%

Open postings that mention RL / post-training

How to learn it

  • RLHF book — Nathan Lambert free

    Post-training as practised on language models — reward models, PPO, DPO, verifiable rewards — by a researcher who post-trains open models, free to read online.

  • TRL documentation — Hugging Face free

    Runnable GRPO and DPO trainers, so the book's methods become an experiment on a laptop-sized model.

Project: Preference- or reward-tune a small model on a verifiable task — GRPO on arithmetic or unit-test-checked code with a small open model; plot reward against a held-out eval and explain where it starts to game the reward.
A posting counts when its description mentions the skill in what the role does or asks for — the company's description of itself is left out. Shares are over technical postings only; the tags are checked against the postings every month and a skill below90% precision is not shown. How we count.