Senior Software Development Engineer in Test (SDET) - AI Cluster
Cerebras · Sunnyvale, CA
About this role
## Senior Software Development Engineer in Test (SDET) — AI Cluster **About Cerebras** Cerebras Systems builds the world’s largest AI chip—56× larger than GPUs. This architecture enables industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services), transforming AI application experiences with real-time iteration and agentic computation. In our AI infrastructure organization, we focus on simplifying large hardware deployments with a “push button” experience and a single pane of glass for observability/monitoring, plus software capabilities built for resiliency. --- ## Responsibilities - Innovate and execute tests on cutting-edge AI infrastructure. - Define optimized test strategies and methodologies. - Adapt quickly to a fast-moving ML ecosystem and bring expertise to new technologies. - Deeply understand large-scale distributed ML training and inference. - Break complex distributed systems into smaller components that can be unit tested. - **Automation-first**: aim for **100% automated tests** across cluster features, including high availability, failure scenarios, performance, stress, and security. - Champion cluster security and reliability (targeting **99.9999% uptime**) and improve ease of use with observability. - Test cluster components, including (but not limited to): **Kubernetes**, **Prometheus**, and **Grafana**. - Test hardware-related components such as ML wafer-scale accelerators, CPU runtime nodes, high-speed interconnects, and high-speed data transfer paths. --- ## Qualifications - Bachelor’s or master’s degree in engineering (CS, Electrical, AI, Data Science, or related). - **5+ years** of experience testing enterprise software, distributed systems, and/or datacenter hardware/software. - Strong coding skills in **Python, Go (Golang), and/or C/C++**. - Strong debugging skills for large distributed systems, hardware, and software. - Experience with tools such as **pdb, gdb, strace**, and network monitoring. - Strong understanding of operating systems internals (memory management, filesystem behavior, security, performance). - Strong understanding of datacenter layout and device performance characteristics (servers, memory, BIOS, PCIe, networking, storage). - Experience with cloud technologies such as **AWS**, **Kubernetes**, and **Docker**. - Monitoring tools like **Grafana** and **Prometheus** are a plus. - Understanding/experience with ML model training and inference is a plus. - Understanding ML hardware accelerators (GPU/custom accelerator ASIC) is a plus. --- ## Why Join Cerebras - Build a breakthrough AI platform beyond GPU constraints. - Publish and open-source cutting-edge AI research. - Work on one of the fastest AI supercomputers in the world. - Job stability with startup vitality. - A simple, non-corporate culture that respects individual beliefs. Learn more: https://www.cerebras.ai/join-us Equal Opportunity Employer — Cerebras is committed to building an equal and diverse environment. Privacy/CCPA: https://www.cerebras.net/privacy/
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.