Open roles
146 postings · verified 2026-10-05 · a repost appears here once per listing but counts once in the statistics (how we count)
- Senior Solution EngineerLambda · San Francisco Office (Second St) | San Jose Office (First St) | Bellevue Office · FDE / Applied · $226K–$355K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonGoC++
- LLM Inference Engineer (Mid, Sr, Staff)Hippocratic AI · Menlo Park, CA · Inference / Perf · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTPythonC++
- Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)Hippocratic AI · Menlo Park, CA · Research Scientist · first seen 2026-09-29EvalsLLMsFine-tuningRL / post-trainingPyTorchvLLM / SGLang / TensorRTDistributed trainingPython
- Manager, Field EngineeringFireworks AI · San Mateo | New York · FDE / Applied · $270K–$310K base · first seen 2026-09-29LLMsFine-tuningvLLM / SGLang / TensorRT
- Member of Technical Staff, LLM InfrastructureFireworks AI · San Mateo · Infra / Hardware · $175K–$220K base · first seen 2026-09-29LLMsvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonGo
- AI Field Engineer - AI NativesFireworks AI · San Mateo | New York · FDE / Applied · $200K–$260K base · first seen 2026-09-29LLMsFine-tuningRL / post-trainingvLLM / SGLang / TensorRTKubernetesPython
- AI Field Engineer - EnterpriseFireworks AI · San Mateo | New York | United States, Remote · FDE / Applied · $200K–$260K base · first seen 2026-09-29EvalsLLMsFine-tuningRL / post-trainingvLLM / SGLang / TensorRTKubernetesPython
- AI Field Engineer, EMEAFireworks AI · London · FDE / Applied · first seen 2026-09-29EvalsLLMsFine-tuningRL / post-trainingvLLM / SGLang / TensorRTKubernetesPython
- AI Field Engineer - Strategic PartnershipsFireworks AI · San Mateo | New York · GTM · $200K–$260K base · first seen 2026-09-29EvalsLLMsFine-tuningvLLM / SGLang / TensorRTKubernetesPython
- AI Field Engineer, SingaporeFireworks AI · Singapore · FDE / Applied · first seen 2026-09-29EvalsLLMsFine-tuningRL / post-trainingvLLM / SGLang / TensorRTKubernetesPython
- Inference InternEtched · San Jose · Inference / Perf · first seen 2026-09-29PyTorchJAXvLLM / SGLang / TensorRTPythonRustC++
- Research Engineer - InferenceElevenLabs · United Kingdom | United States | Poland | Bulgaria · Remote · Research Engineer · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRT
- ML Ops Infrastructure EngineerDeepgram · USA - Remote · Remote · Infra / Hardware · $160K–$220K base · first seen 2026-09-29EvalsLLMsvLLM / SGLang / TensorRTKubernetesPython
- Senior Technical Program Manager (Engineering) - AI Tooling & SystemsDeepgram · USA - Remote · Remote · Program / Ops · $152K–$208K base · first seen 2026-09-29EvalsRAG / retrievalLLMsPromptingPyTorchCUDA / TritonvLLM / SGLang / TensorRT
- Senior Staff Applied AI Inference EngineerCrusoe · San Francisco, CA - US · FDE / Applied · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTKubernetesPythonC++
- Engineering Manager, AI PlatformCrusoe · San Francisco, CA - US | Sunnyvale, CA - US · Infra / Hardware · first seen 2026-09-29LLMsvLLM / SGLang / TensorRTKubernetesPythonRust
- Senior Manager, Engineering - AI InferenceCrusoe · San Francisco, CA - US · Inference / Perf · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTKubernetesPythonC++
- Staff Applied AI Inference EngineerCrusoe · San Francisco, CA - US · FDE / Applied · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTKubernetesPythonC++
- Product Manager, Managed NorthCohere · Toronto | Canada | United States · Remote · Product / Design · $160K–$285K base · first seen 2026-09-29vLLM / SGLang / TensorRT
- Audio Inference Engineer, Model EfficiencyCohere · New York | San Francisco | Toronto | Montreal · Remote · Inference / Perf · $205K–$380K base · first seen 2026-09-29LLMsPyTorchvLLM / SGLang / TensorRTPythonC++
- Senior ML Systems Engineer, Frameworks & ToolingCohere · London | San Francisco | New York | Paris | Toronto | Montreal · Remote · Software Eng · $205K–$380K base · first seen 2026-09-29LLMsPyTorchJAXCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetes
- Senior Member of Technical Staff, Synthetic DataCohere · Toronto | London | New York | Paris | Montreal · Remote · Research Engineer · $205K–$380K base · first seen 2026-09-29LLMsvLLM / SGLang / TensorRTPython
- Member of Technical Staff, Model EfficiencyCohere · New York | San Francisco | Toronto | Montreal · Remote · Research Engineer · $205K–$380K base · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTPythonRustGoC++
- Staff Software Engineer, GPU InferenceCerebras · Toronto, CAN | Sunnyvale, CA · Inference / Perf · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonC++
- Staff Software Engineer, Inference APICerebras · Toronto, CAN · Inference / Perf · first seen 2026-09-29LLMsPyTorchvLLM / SGLang / TensorRTKubernetesPythonRustGoC++
- Senior Product Manager, AI ModelsCerebras · Sunnyvale, CA | Remote (US) · Product / Design · first seen 2026-09-29EvalsLLMsPromptingPyTorchvLLM / SGLang / TensorRTPython
- Software Engineer, GPU InferenceCerebras · United States and Canada · Inference / Perf · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonC++
- Inference EngineerCartesia · *HQ - San Francisco, CA · Inference / Perf · $180K–$250K base · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRT
- Member of Technical Staff - Model Serving / API Backend EngineerBlack Forest Labs · San Francisco (United States) · Software Eng · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRTKubernetesPython
- Software Engineer - Model PerformanceBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $180K–$360K base · first seen 2026-09-29LLMsPyTorchCUDA / TritonvLLM / SGLang / TensorRTPythonC++
- Software Engineer - Baseten Inference StackBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $180K–$360K base · first seen 2026-09-29vLLM / SGLang / TensorRTKubernetes
- Solutions ArchitectBaseten · San Francisco | New York · FDE / Applied · $165K–$330K base · first seen 2026-09-29LLMsvLLM / SGLang / TensorRT
- Software Engineer - Voice AI (Inference Runtime)Baseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $165K–$330K base · first seen 2026-09-29vLLM / SGLang / TensorRTKubernetesPython
- Technical Program Manager, Model PerformanceBaseten · San Francisco | Remote | Toronto | New York | Montreal | Seattle · Program / Ops · $165K–$330K base · first seen 2026-09-29vLLM / SGLang / TensorRT
- Product Manager, Inference PlatformBaseten · San Francisco | New York · Product / Design · $235K–$335K base · first seen 2026-09-29vLLM / SGLang / TensorRTKubernetes
- Software Engineer - Model ProductsBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Software Eng · $180K–$360K base · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRTKubernetes
- Software Engineer - GPU Networking & Distributed SystemsBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $165K–$330K base · first seen 2026-09-29vLLM / SGLang / TensorRTDistributed trainingPythonRustC++
- Software Engineer - Training ProductBaseten · San Francisco | New York · Software Eng · $165K–$330K base · first seen 2026-09-29Fine-tuningRL / post-trainingPyTorchvLLM / SGLang / TensorRTDistributed trainingKubernetes
- Forward Deployed Engineer (Training)Baseten · San Francisco | New York · FDE / Applied · $200K–$400K base · first seen 2026-09-29EvalsLLMsRL / post-trainingPyTorchJAXvLLM / SGLang / TensorRTKubernetes
- Senior Software Engineer - ML InfrastructureApplied Intuition · Sunnyvale · Infra / Hardware · $215K–$285K base · first seen 2026-09-29PyTorchvLLM / SGLang / TensorRT
- ML Runtime Optimization EngineerApplied Intuition · Sunnyvale · Software Eng · $217K–$318K base · first seen 2026-09-29PyTorchJAXCUDA / TritonvLLM / SGLang / TensorRT
- AI Performance EngineerApplied Intuition · Sunnyvale · Inference / Perf · $215K–$285K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingPythonC++
- Distributed LLM Inference EngineerAnyscale · San Francisco | Palo Alto · Inference / Perf · $170K–$245K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRT
- Machine Learning Engineer, Customer EngineeringAnyscale · San Francisco · FDE / Applied · $170K–$199K base · first seen 2026-09-29LLMsvLLM / SGLang / TensorRTKubernetes
- Member of Technical Staff, Machine Learning InfrastructureAbridge · SF Office · Infra / Hardware · $221K–$260K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetes
- AI Researcher1X · San Carlos, CA · Research Scientist · $250K–$350K base · first seen 2026-09-29EvalsLLMsPyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingPython