CronJobs

backend jobs

Software Engineer, GPU Inference

Cerebras · Sunnyvale, CA

onsiteseniorPosted Nov 25, 2025C++PythonvLLMPyTorchROCmHIPKubernetesLinux

Apply on the employer site

About this role

## Software Engineer, GPU Inference ### About Cerebras Cerebras Systems builds the world’s largest AI chip—56× larger than GPUs—enabling industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services). Cerebras partners with leading model labs, global enterprises, and AI-native startups, including a multi-year partnership with OpenAI to deploy 750MW of scale. ### About the Role Cerebras is building a new generation of **disaggregated AI inference systems** that combine **GPU-accelerated prefill** with **ultra-fast decode** on the **Cerebras Wafer-Scale Engine**. You’ll help **productionize and optimize the GPU serving stack**, working across: - Custom inference APIs - **vLLM** serving runtime - **AMD ROCm** software stack - Rack-scale **AMD GPU infrastructure** This is a hands-on role focused on making the serving path **reliable, numerically correct, observable, and exceptionally performant**—improving **time to first token, throughput, tail latency, and capacity efficiency**. ### Responsibilities - **Productionize the GPU inference stack**: design, build, deploy, and maintain the complete GPU prefill path (API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, rack-scale infrastructure). - **Own GPU operational readiness**: deployment/upgrade/rollback, health checks, capacity management, failure recovery; automation for driver/firmware/runtime/model/container compatibility. - **Drive reliability in production**: define SLOs/indicators; improve fault isolation, graceful degradation, automated recovery, incident response, and remediation. - **Improve inference performance**: optimize time to first token, throughput, tokens/sec/GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity. - **Optimize model-serving behavior**: scheduling, continuous batching, prefix caching, KV-cache management, tensor/expert parallelism, request admission, quantization, graph execution, and distributed communication. - **Debug across system layers**: diagnose failures/performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective libraries, kernels, drivers/firmware, networking, and hardware. - **Ensure numerical correctness**: validation/regression infrastructure for quality, numerical accuracy, precision changes, quantization, determinism, and compatibility. - **Build performance + correctness infrastructure**: benchmarks, workload replay, profiling automation, release qualification, dashboards, and regression gates. ### Minimum Qualifications - **5+ years** software engineering experience with ownership of complex production systems. - Experience building/operating/optimizing **production inference systems** for LLMs or similarly demanding GPU workloads. - Strong **C++ and Python** skills (multithreading, concurrency, memory management, performance-sensitive development). - Hands-on experience with high-performance model-serving frameworks (e.g., **vLLM**, SGLang, TensorRT-LLM, Triton, or equivalent). - Strong understanding of **GPU execution/performance** (async execution, memory movement, synchronization, kernel launches, profiling). - Experience debugging **distributed systems** across layers (not treating the serving framework/runtime as a black box). - Experience with **Linux**, containers, **Kubernetes** (or similar), observability, CI/CD, and operating latency-sensitive services. - Ability to design rigorous benchmarks a

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord