Skill · AI Hiring Index · as of 2026-10-05
Evals in AI job postings
15% of technical postings at the AI companies we track mention Evals (738 postings). AI / ML Engineer postings ask for it most: 51%.
Technical postings that mention it
15%
738 postings
Which roles ask for it
Share of each technical role family's open postings that mention Evals.
| Role family | Share, as a bar | Share |
|---|---|---|
| AI / ML Engineer | 51% | |
| Research Engineer | 47% | |
| Research Scientist | 45% | |
| Evals / Data | 44% | |
| FDE / Applied | 21% | |
| Safety / Policy | 21% | |
| Software Eng | 7% | |
| Inference / Perf | 6% | |
| Infra / Hardware | 3% | |
| Security | 1% |
Companies asking for it most
| Company | Postings | Of its technical postings |
|---|---|---|
| OpenAI | 105 | 22% |
| Anthropic | 77 | 23% |
| Sierra | 54 | 73% |
| Scale AI | 42 | 40% |
| Decagon | 29 | 38% |
| Databricks | 29 | 5% |
| Cohere | 24 | 31% |
| LangChain | 22 | 43% |
| Harvey | 20 | 20% |
| Mistral AI | 20 | 17% |
How to learn it
- AI Evals: Everything You Need to Know — Hamel Husain and Shreya Shankar free
Start from error analysis on real traces, not from a metric; then judges you check against people — the practice behind 'evaluation frameworks' in these postings.
Project: An eval harness that gates a change — Label 100 real outputs by hand, build an LLM judge and measure its agreement with you, and make CI fail when a prompt or model change drops the score.