CronJobs

devops-sre jobs

Staff GPU Inference SDET

Cerebras · Sunnyvale, CA

hybridstaffPosted Sep 9, 2026PythonKubernetesPrometheusGrafanaPyTorchNCCLSlurmRay

Apply on the employer site

About this role

## About the Role Cerebras Systems builds the world’s largest AI chip (56x larger than GPUs), enabling industry-leading training and inference speeds—over 10x faster than GPU-based hyperscale cloud inference services. As a **Staff GPU Inference SDET**, you’ll be the founding **quality, reliability, and validation lead** for a new GPU Inference Development team. You’ll design, build, and scale the end-to-end **release qualification** and **automated test ecosystem** for the GPU inference stack and rack-scale accelerated compute fleets. You’ll own automated test suites for **multi-node GPU cluster bring-up**, **prefill worker optimizations**, validation of **open-source and custom serving engines**, and ensuring **numerical correctness** and **performance stability** under real-world streaming workloads. ## What You’ll Do - **Build GPU Release Qualification Systems**: Automated test frameworks, regression gates, and release qualification pipelines across the GPU inference stack (API services, model-serving workers, container runtimes, serving engines, driver stacks, firmware). - **Inference Serving & Workload Validation**: Benchmark and stress-test distributed LLM serving frameworks (prefill vs. decode performance, continuous batching, prefix caching, KV-cache efficiency, tensor/expert parallelism). - **Performance & Performance Modeling Verification**: Workload replay and benchmarking tools to validate GPU performance models. - Track metrics such as **Time-to-First-Token (TTFT)**, **Inter-Token Latency (ITL)**, throughput, **P99 tail latency**, and capacity efficiency. - **Numerical Correctness & Quality Gates**: Validation infrastructure for accuracy, precision stability (**FP16/FP8/quantization**), determinism, and output correctness across software updates, kernel fusions, and hardware revisions. - **Fault Injection & Fleet Resilience**: Chaos engineering and fault-injection suites (node failures, network degradation, GPU memory leaks, driver/firmware mismatches, automated recovery). - **Observability & CI/CD Integration**: Integrate automated test pipelines with telemetry tools (e.g., **Prometheus, Grafana**) for repeatable engineering gates and continuous performance monitoring. ## Requirements - **8+ years** of software engineering experience as an **SDET**, **Infrastructure Quality Lead**, or **Systems Test Engineer**. - **GPU & Cluster Infrastructure Expertise**: Hands-on experience bringing up, provisioning, and validating **multi-node GPU clusters** (NVIDIA or AMD) in public cloud or enterprise data centers. - **Inference Stack Knowledge**: Deep understanding of LLM serving engines and distributed runtimes (prefill vs. decode, KV-cache management, dynamic batching). - **Automation & Scripting**: Expert **Python** skills; experience building custom test automation frameworks, diagnostic tooling, and CI/CD integration. - **Orchestration & Networking**: Proficiency with **Kubernetes, Slurm, Ray** and high-performance interconnects (**InfiniBand, RoCE, NCCL**). - **Failure Analysis & Debugging**: Root-cause analysis across software/hardware boundaries; stress testing and node failure simulation in distributed systems. ## Nice to Haves - Direct experience with **AMD (ROCm / HIP)** or **NVIDIA** software stacks. - Experience building workload replay tools, ML evaluation pipelines, or **MLPerf Inference** benchmark suites. - Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or **C++**.

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord