CronJobs

devops-sre jobs

Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras · Sunnyvale, CA

hybridseniorPosted Jul 15, 2026PythonKubernetesAWSPrometheusGrafanaGolangC/C++Docker

Apply on the employer site

About this role

## Senior Software Development Engineer in Test (SDET) — AI Cluster **About Cerebras** Cerebras Systems builds the world’s largest AI chip—56× larger than GPUs. This architecture enables industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services), transforming AI application experiences with real-time iteration and agentic computation. In our AI infrastructure organization, we focus on simplifying large hardware deployments with a “push button” experience and a single pane of glass for observability/monitoring, plus software capabilities built for resiliency. --- ## Responsibilities - Innovate and execute tests on cutting-edge AI infrastructure. - Define optimized test strategies and methodologies. - Adapt quickly to a fast-moving ML ecosystem and bring expertise to new technologies. - Deeply understand large-scale distributed ML training and inference. - Break complex distributed systems into smaller components that can be unit tested. - **Automation-first**: aim for **100% automated tests** across cluster features, including high availability, failure scenarios, performance, stress, and security. - Champion cluster security and reliability (targeting **99.9999% uptime**) and improve ease of use with observability. - Test cluster components, including (but not limited to): **Kubernetes**, **Prometheus**, and **Grafana**. - Test hardware-related components such as ML wafer-scale accelerators, CPU runtime nodes, high-speed interconnects, and high-speed data transfer paths. --- ## Qualifications - Bachelor’s or master’s degree in engineering (CS, Electrical, AI, Data Science, or related). - **5+ years** of experience testing enterprise software, distributed systems, and/or datacenter hardware/software. - Strong coding skills in **Python, Go (Golang), and/or C/C++**. - Strong debugging skills for large distributed systems, hardware, and software. - Experience with tools such as **pdb, gdb, strace**, and network monitoring. - Strong understanding of operating systems internals (memory management, filesystem behavior, security, performance). - Strong understanding of datacenter layout and device performance characteristics (servers, memory, BIOS, PCIe, networking, storage). - Experience with cloud technologies such as **AWS**, **Kubernetes**, and **Docker**. - Monitoring tools like **Grafana** and **Prometheus** are a plus. - Understanding/experience with ML model training and inference is a plus. - Understanding ML hardware accelerators (GPU/custom accelerator ASIC) is a plus. --- ## Why Join Cerebras - Build a breakthrough AI platform beyond GPU constraints. - Publish and open-source cutting-edge AI research. - Work on one of the fastest AI supercomputers in the world. - Job stability with startup vitality. - A simple, non-corporate culture that respects individual beliefs. Learn more: https://www.cerebras.ai/join-us Equal Opportunity Employer — Cerebras is committed to building an equal and diverse environment. Privacy/CCPA: https://www.cerebras.net/privacy/

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord