Staff GPU Inference SDET
Cerebras · Sunnyvale, CA
About this role
## About the Role Cerebras Systems builds the world’s largest AI chip (56x larger than GPUs), enabling industry-leading training and inference speeds—over 10x faster than GPU-based hyperscale cloud inference services. As a **Staff GPU Inference SDET**, you’ll be the founding **quality, reliability, and validation lead** for a new GPU Inference Development team. You’ll design, build, and scale the end-to-end **release qualification** and **automated test ecosystem** for the GPU inference stack and rack-scale accelerated compute fleets. You’ll own automated test suites for **multi-node GPU cluster bring-up**, **prefill worker optimizations**, validation of **open-source and custom serving engines**, and ensuring **numerical correctness** and **performance stability** under real-world streaming workloads. ## What You’ll Do - **Build GPU Release Qualification Systems**: Automated test frameworks, regression gates, and release qualification pipelines across the GPU inference stack (API services, model-serving workers, container runtimes, serving engines, driver stacks, firmware). - **Inference Serving & Workload Validation**: Benchmark and stress-test distributed LLM serving frameworks (prefill vs. decode performance, continuous batching, prefix caching, KV-cache efficiency, tensor/expert parallelism). - **Performance & Performance Modeling Verification**: Workload replay and benchmarking tools to validate GPU performance models. - Track metrics such as **Time-to-First-Token (TTFT)**, **Inter-Token Latency (ITL)**, throughput, **P99 tail latency**, and capacity efficiency. - **Numerical Correctness & Quality Gates**: Validation infrastructure for accuracy, precision stability (**FP16/FP8/quantization**), determinism, and output correctness across software updates, kernel fusions, and hardware revisions. - **Fault Injection & Fleet Resilience**: Chaos engineering and fault-injection suites (node failures, network degradation, GPU memory leaks, driver/firmware mismatches, automated recovery). - **Observability & CI/CD Integration**: Integrate automated test pipelines with telemetry tools (e.g., **Prometheus, Grafana**) for repeatable engineering gates and continuous performance monitoring. ## Requirements - **8+ years** of software engineering experience as an **SDET**, **Infrastructure Quality Lead**, or **Systems Test Engineer**. - **GPU & Cluster Infrastructure Expertise**: Hands-on experience bringing up, provisioning, and validating **multi-node GPU clusters** (NVIDIA or AMD) in public cloud or enterprise data centers. - **Inference Stack Knowledge**: Deep understanding of LLM serving engines and distributed runtimes (prefill vs. decode, KV-cache management, dynamic batching). - **Automation & Scripting**: Expert **Python** skills; experience building custom test automation frameworks, diagnostic tooling, and CI/CD integration. - **Orchestration & Networking**: Proficiency with **Kubernetes, Slurm, Ray** and high-performance interconnects (**InfiniBand, RoCE, NCCL**). - **Failure Analysis & Debugging**: Root-cause analysis across software/hardware boundaries; stress testing and node failure simulation in distributed systems. ## Nice to Haves - Direct experience with **AMD (ROCm / HIP)** or **NVIDIA** software stacks. - Experience building workload replay tools, ML evaluation pipelines, or **MLPerf Inference** benchmark suites. - Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or **C++**.
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.