Skill · AI Hiring Index · as of 2026-10-05
vLLM / SGLang / TensorRT in AI job postings
3% of technical postings at the AI companies we track mention vLLM / SGLang / TensorRT (133 postings). Inference / Perf postings ask for it most: 22%.
Technical postings that mention it
3%
133 postings
Which roles ask for it
Share of each technical role family's open postings that mention vLLM / SGLang / TensorRT.
| Role family | Share, as a bar | Share |
|---|---|---|
| Inference / Perf | 22% | |
| AI / ML Engineer | 5% | |
| Research Engineer | 4% | |
| Research Scientist | 4% | |
| FDE / Applied | 4% | |
| Software Eng | 1% | |
| Infra / Hardware | 1% | |
| Evals / Data | 0% | |
| Safety / Policy | 0% | |
| Security | 0% |
Companies asking for it most
| Company | Postings | Of its technical postings |
|---|---|---|
| Nebius | 24 | 11% |
| CoreWeave | 17 | 10% |
| Baseten | 8 | 19% |
| SambaNova | 7 | 17% |
| Together AI | 7 | 13% |
| Fireworks AI | 6 | 19% |
| Prime Intellect | 5 | 24% |
| Cohere | 4 | 5% |
| OpenAI | 4 | 1% |
| Cerebras | 4 | 4% |
How to learn it
- LLM Inference Handbook — Modular free
Latency, throughput and cost of serving a model — batching, KV cache, quantization, which engine when — as one practical handbook rather than twenty blog posts.
Project: Serve a small open model and measure it — Serve a small open model with vLLM, then measure time to first token and tokens per second at 1, 8 and 32 concurrent requests, with and without quantization.