← all roles

Inference & Model Serving Jobs

Running models in production — inference engines, model serving, and latency/throughput optimization (vLLM, TensorRT and similar). 47 open now, refreshed daily.

open roles
47
companies
17
list salary
27 · $139K–$625K
visa mention
9
remote
4

Observed across current open postings, refreshed daily — not a survey. Salary band is drawn only from roles that publish a range. Salary breakdown →

Inference and model-serving roles own the production side: getting trained models to answer fast and cheaply under real traffic. That means serving engines and runtimes (vLLM, TensorRT-LLM and the like), continuous batching and KV-cache strategy, quantization, and the latency/throughput trade-offs that decide unit economics for anyone shipping an LLM product. They concentrate at the labs and inference-platform startups whose revenue is literally tokens-per-second — so the work rewards people who reason fluently about both model internals and the systems that run them.

Hiring most for this specialty: Together AI 9 · Anthropic 8 · CoreWeave 6 · Databricks 5 · Nebius 3 · Baseten 2 · see all who's hiring →

filter
view
47 roles · refreshed 2026-09-03 10:46 UTC