Software Engineer, Cluster Deployment
Cerebras · Sunnyvale, CA
About this role
## Software Engineer, Cluster Deployment ### About the Role Cerebras builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (over 10x faster than GPU-based hyperscale cloud inference). You’ll help build and operate the software systems that deploy, validate, and manage AI compute clusters across data centers worldwide. This includes turning complex bare-metal infrastructure into repeatable, automated deployment flows covering: - Server provisioning - Network configuration - Kubernetes bring-up - Health validation - Operational handoff As a Software Engineer on the **Cluster Deployment Automation** team, you’ll build “pushbutton” tooling to make large-scale cluster deployments faster, safer, and more reproducible. ### Responsibilities - Develop and maintain automation for deployment workflows (provisioning, configuration, validation, operational handoff) - Convert manual deployment steps into tested, repeatable workflows - Participate in hands-on cluster deployments to build practical debugging and operational expertise - Troubleshoot issues across Linux, bare-metal servers, networking, storage, Kubernetes, and connectivity - Contribute to infrastructure-as-code and GitOps workflows (e.g., Terraform, Ansible, PRs, code review) - Add health checks, observability, dashboards, and validation logic to improve reliability - Partner with networking, infrastructure, security, and operations teams to deliver secure, reproducible deployments ### Basic Qualifications - 2+ years of mid- to large-scale data center deployment experience - Strong fundamentals in **Python** and **Bash** (ability to write scripts and small programs) - Basic **Linux** experience (CLI, processes, filesystems, disk troubleshooting) - Working knowledge of **Git** (branching, commits, PRs, code review) - CS/ECE or related technical degree, or equivalent practical experience - Curiosity, strong problem-solving, and willingness to work hands-on with real infrastructure ### Preferred Qualifications - Networking fundamentals (VLANs, routing); exposure to BGP, switch configuration, or Arista EOS automation is a plus - Kubernetes experience or familiarity - Infrastructure-as-code and GitOps experience (Terraform, Ansible, PR-based change control) - Bare-metal provisioning concepts (PXE, DHCP, iPXE, Redfish, IPMI, BMC management) - Observability experience (Prometheus or Grafana) - API design and client-server architecture - Automation side projects or open-source contributions ### Why Join Cerebras - Build a breakthrough AI platform beyond GPU constraints - Publish and open source cutting-edge AI research - Work on one of the fastest AI supercomputers in the world - Startup vitality with job stability - A simple, non-corporate work culture **Apply today** and help drive groundbreaking advancements in AI! *Equal Opportunity Employer.*
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.