Cluster Operations Software Engineer
Cerebras · Sunnyvale, CA
About this role
## Cluster Operations Software Engineer ### About Cerebras Cerebras Systems builds the world’s largest AI chip—**the Wafer-Scale Engine (WSE)**—enabling training and inference at speeds **over 10× faster than GPU-based hyperscale cloud inference**. Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy **750MW** of scale. ### The Role As an **AI Cluster Operations Engineer**, you’ll manage and operate cutting-edge machine learning compute clusters that harness the WSE. You’ll ensure **health, performance, and availability** of infrastructure, maximize compute capacity, and support growing AI initiatives. ### Responsibilities - Deploy, configure, and debug **container-based services** using **Docker** - Build and own cluster operations software (e.g., **monitoring**, **workflow automation**, **operational dashboards**, **reliability tooling**) - Collaborate with cross-functional teams to translate operational needs into scalable O&M products and platform capabilities - Develop **APIs**, automation services, and integrations to improve operational visibility, incident response, and fleet management - Manage and operate multiple advanced AI compute infrastructure clusters - Monitor cluster health and proactively identify and resolve issues - Maximize compute capacity via optimization and efficient resource allocation - Provide **24/7 monitoring and support**, including hands-on troubleshooting - Handle engineering escalations and collaborate to resolve complex technical challenges - Stay current with advancements in AI compute infrastructure and related technologies ### Skills & Requirements - **6–8 years** managing/operating complex compute infrastructure (ML or HPC preferred) - Proficiency in **Python** and **Go**; experience building operational platforms, automation, and reliability tooling - **Distributed systems** expertise (must) - Deep understanding of **Linux** and command-line tools - Extensive knowledge of **Docker** and container orchestration (e.g., **Kubernetes / k8s**) - Strong troubleshooting skills for complex technical issues - Experience with **monitoring and alerting** systems - Proven ability to own and drive challenges to completion - Excellent communication and collaboration skills - Ability to work in a fast-paced environment - Willingness to participate in a **24/7 on-call rotation** ### Preferred - Experience operating/managing large-scale AI clusters - Knowledge of networking technologies (e.g., **Ethernet, RoCE, TCP/IP**) - Knowledge of cloud platforms (**AWS, GCP, Azure**) ### Location - **SF Bay Area** - **Toronto, Canada** - **Bangalore, India** ### Why Join Cerebras People who are serious about software make their own hardware. Join a team building a breakthrough AI platform beyond GPU constraints, publishing/open-sourcing cutting-edge research, and working on one of the fastest AI supercomputers in the world. Apply today and help drive groundbreaking advancements in AI!
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.