Software Engineer, GPU Inference
Cerebras · Sunnyvale, CA
About this role
## Software Engineer, GPU Inference ### About Cerebras Cerebras Systems builds the world’s largest AI chip—56× larger than GPUs—enabling industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services). Cerebras partners with leading model labs, global enterprises, and AI-native startups, including a multi-year partnership with OpenAI to deploy 750MW of scale. ### About the Role Cerebras is building a new generation of **disaggregated AI inference systems** that combine **GPU-accelerated prefill** with **ultra-fast decode** on the **Cerebras Wafer-Scale Engine**. You’ll help **productionize and optimize the GPU serving stack**, working across: - Custom inference APIs - **vLLM** serving runtime - **AMD ROCm** software stack - Rack-scale **AMD GPU infrastructure** This is a hands-on role focused on making the serving path **reliable, numerically correct, observable, and exceptionally performant**—improving **time to first token, throughput, tail latency, and capacity efficiency**. ### Responsibilities - **Productionize the GPU inference stack**: design, build, deploy, and maintain the complete GPU prefill path (API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, rack-scale infrastructure). - **Own GPU operational readiness**: deployment/upgrade/rollback, health checks, capacity management, failure recovery; automation for driver/firmware/runtime/model/container compatibility. - **Drive reliability in production**: define SLOs/indicators; improve fault isolation, graceful degradation, automated recovery, incident response, and remediation. - **Improve inference performance**: optimize time to first token, throughput, tokens/sec/GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity. - **Optimize model-serving behavior**: scheduling, continuous batching, prefix caching, KV-cache management, tensor/expert parallelism, request admission, quantization, graph execution, and distributed communication. - **Debug across system layers**: diagnose failures/performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective libraries, kernels, drivers/firmware, networking, and hardware. - **Ensure numerical correctness**: validation/regression infrastructure for quality, numerical accuracy, precision changes, quantization, determinism, and compatibility. - **Build performance + correctness infrastructure**: benchmarks, workload replay, profiling automation, release qualification, dashboards, and regression gates. ### Minimum Qualifications - **5+ years** software engineering experience with ownership of complex production systems. - Experience building/operating/optimizing **production inference systems** for LLMs or similarly demanding GPU workloads. - Strong **C++ and Python** skills (multithreading, concurrency, memory management, performance-sensitive development). - Hands-on experience with high-performance model-serving frameworks (e.g., **vLLM**, SGLang, TensorRT-LLM, Triton, or equivalent). - Strong understanding of **GPU execution/performance** (async execution, memory movement, synchronization, kernel launches, profiling). - Experience debugging **distributed systems** across layers (not treating the serving framework/runtime as a black box). - Experience with **Linux**, containers, **Kubernetes** (or similar), observability, CI/CD, and operating latency-sensitive services. - Ability to design rigorous benchmarks a
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.