Open roles
177 postings · verified 2026-10-05 · a repost appears here once per listing but counts once in the statistics (how we count)
- Senior Manager, Engineering - AI InferenceCrusoe · San Francisco, CA - US · Inference / Perf · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTKubernetesPythonC++
- Staff Applied AI Inference EngineerCrusoe · San Francisco, CA - US · FDE / Applied · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTKubernetesPythonC++
- Member of Technical Staff, Training Performance EngineerCohere · London | New York | Paris | Toronto | Montreal · Inference / Perf · $205K–$380K base · first seen 2026-09-29PyTorchJAXCUDA / TritonDistributed trainingPython
- Senior ML Systems Engineer, Frameworks & ToolingCohere · London | San Francisco | New York | Paris | Toronto | Montreal · Remote · Software Eng · $205K–$380K base · first seen 2026-09-29LLMsPyTorchJAXCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetes
- Senior Member of Technical Staff, Multimodal AICohere · San Francisco | New York | Paris | Toronto | Montreal · Remote · Research Engineer · $205K–$380K base · first seen 2026-09-29EvalsLLMsPyTorchJAXCUDA / TritonDistributed trainingPython
- Machine Learning Intern/Co-op (Winter 2027)Cohere · Canada | Europe | United States | United Kingdom · Remote · Software Eng · first seen 2026-09-29LLMsJAXCUDA / TritonDistributed trainingPython
- Member of Technical Staff, ModelingCohere · London | San Francisco | New York | Paris | Toronto | Montreal · Remote · Research Engineer · $205K–$380K base · first seen 2026-09-29LLMsJAXCUDA / TritonDistributed trainingPython
- Member of Technical Staff, Model EfficiencyCohere · New York | San Francisco | Toronto | Montreal · Remote · Research Engineer · $205K–$380K base · first seen 2026-09-29LLMsCUDA / TritonvLLM / SGLang / TensorRTPythonRustGoC++
- ML Algorithm Mapping and Performance Engineer, Core MLCerebras · Sunnyvale, CA · Inference / Perf · first seen 2026-09-29PyTorchJAXCUDA / TritonDistributed trainingPythonC++
- ML Runtime and Kernel Engineer - Core MLCerebras · Sunnyvale, CA | Toronto, CAN · Inference / Perf · first seen 2026-09-29LLMsPyTorchJAXCUDA / TritonPythonC++
- Staff Software Engineer, GPU InferenceCerebras · Toronto, CAN | Sunnyvale, CA · Inference / Perf · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonC++
- Principal ML InvestigatorCerebras · Sunnyvale, CA · Other · first seen 2026-09-29LLMsRL / post-trainingCUDA / TritonDistributed trainingPhD
- Software Engineer, GPU InferenceCerebras · United States and Canada · Inference / Perf · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetesPythonC++
- Inference EngineerCartesia · *HQ - San Francisco, CA · Inference / Perf · $180K–$250K base · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRT
- Member of Technical Staff - Image / Video GenerationBlack Forest Labs · Freiburg (Germany) · Research Engineer · first seen 2026-09-29LLMsFine-tuningPyTorchCUDA / TritonDistributed training
- Member of Technical Staff - Model Serving / API Backend EngineerBlack Forest Labs · San Francisco (United States) · Software Eng · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRTKubernetesPython
- Member of Technical Staff - Research EngineerBlack Forest Labs · San Francisco (United States) | Freiburg (Germany) · Research Engineer · first seen 2026-09-29LLMsPyTorchCUDA / TritonDistributed training
- Software Engineer - GPU KernelsBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $180K–$360K base · first seen 2026-09-29CUDA / TritonC++
- Software Engineer - Model PerformanceBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Inference / Perf · $180K–$360K base · first seen 2026-09-29LLMsPyTorchCUDA / TritonvLLM / SGLang / TensorRTPythonC++
- Software Engineer - Model ProductsBaseten · San Francisco | Toronto | New York | Montreal | Seattle · Software Eng · $180K–$360K base · first seen 2026-09-29CUDA / TritonvLLM / SGLang / TensorRTKubernetes
- ML Runtime Optimization EngineerApplied Intuition · Sunnyvale · Software Eng · $217K–$318K base · first seen 2026-09-29PyTorchJAXCUDA / TritonvLLM / SGLang / TensorRT
- Software Engineer - E2E AutonomyApplied Intuition · Sunnyvale · Software Eng · $153K–$222K base · first seen 2026-09-29PyTorchCUDA / TritonKubernetesPython
- AI Performance EngineerApplied Intuition · Sunnyvale · Inference / Perf · $215K–$285K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingPythonC++
- Research Engineer - AI/RL InfrastructureApplied Intuition · Sunnyvale · Research Engineer · $126K–$423K base · first seen 2026-09-29EvalsPyTorchCUDA / TritonDistributed trainingKubernetesPhD
- Distributed LLM Inference EngineerAnyscale · San Francisco | Palo Alto · Inference / Perf · $170K–$245K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRT
- Member of Technical Staff, Machine Learning InfrastructureAbridge · SF Office · Infra / Hardware · $221K–$260K base · first seen 2026-09-29PyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingKubernetes
- AI Researcher1X · San Carlos, CA · Research Scientist · $250K–$350K base · first seen 2026-09-29EvalsLLMsPyTorchCUDA / TritonvLLM / SGLang / TensorRTDistributed trainingPython