Skip to main content

Skill · AI Hiring Index · as of 2026-10-05

vLLM / SGLang / TensorRT in AI job postings

3% of technical postings at the AI companies we track mention vLLM / SGLang / TensorRT (133 postings). Inference / Perf postings ask for it most: 22%.

Technical postings that mention it
3%
133 postings

Which roles ask for it

Share of each technical role family's open postings that mention vLLM / SGLang / TensorRT.

Role familyShare, as a barShare
Inference / Perf22%
AI / ML Engineer5%
Research Engineer4%
Research Scientist4%
FDE / Applied4%
Software Eng1%
Infra / Hardware1%
Evals / Data0%
Safety / Policy0%
Security0%

Companies asking for it most

CompanyPostingsOf its technical postings
Nebius2411%
CoreWeave1710%
Baseten819%
SambaNova717%
Together AI713%
Fireworks AI619%
Prime Intellect524%
Cohere45%
OpenAI41%
Cerebras44%

Open postings that mention vLLM / SGLang / TensorRT

How to learn it

  • LLM Inference Handbook — Modular free

    Latency, throughput and cost of serving a model — batching, KV cache, quantization, which engine when — as one practical handbook rather than twenty blog posts.

Project: Serve a small open model and measure it — Serve a small open model with vLLM, then measure time to first token and tokens per second at 1, 8 and 32 concurrent requests, with and without quantization.
A posting counts when its description mentions the skill in what the role does or asks for — the company's description of itself is left out. Shares are over technical postings only; the tags are checked against the postings every month and a skill below90% precision is not shown. How we count.