AI Fleet Platform Software Engineer
Cerebras · Sunnyvale, CA
About this role
## AI Fleet Platform Software Engineer ### About the role Cerebras Systems builds large-scale AI compute infrastructure and operates AI clusters across multiple data centers. As the fleet grows, we need a senior software engineer to build the platforms and tools that help teams monitor fleet health, respond to issues, and keep compute capacity available. You’ll own substantial engineering work across **backend services**, **operational tools**, and **user-facing applications**, partnering with Cluster Operations, infrastructure, inference, security, and data center teams to turn real operational challenges into reliable software at scale. ### Responsibilities - Build and operate software for managing large fleets of AI clusters - Provide operators clear, actionable views of cluster health, capacity, performance, and ongoing issues - Develop services and integrations that bring together data and workflows from multiple infrastructure systems - Automate repetitive work and improve tools for investigating incidents and restoring service - Design systems that remain reliable as the fleet grows, including through component and site failures - Work closely with platform users to understand needs and make practical product/engineering decisions - Lead projects from initial design through production, measure impact, and use operational feedback to improve ### Skills & Requirements - **12+ years** building and operating production distributed systems or large-scale infrastructure software - Strong **Go** skills (regular work with Go, Protocol Buffers, and SQL). **Python** is a plus - Experience building **control planes**, **fleet management systems**, or **operational platforms** - Experience designing **APIs and data models** for multiple consumers, with thoughtful versioning/compatibility - Experience with **asynchronous systems** supporting idempotency, concurrency, and reconciliation while preventing state drift - Strong testing practices, especially for software that can change the state of running production clusters - Experience with **Linux**, **containers**, and **Kubernetes** - Experience with durable workflow execution, event streaming, or high-cardinality time-series telemetry - Sound judgment around **reliability, security, and observability** (including dependency unavailability) - Ability to take an ambiguous problem from design to production and improve using operational feedback ### Preferred Experience - Infrastructure software for thousands of hosts across multiple sites - Incident response, hardware health, or capacity management software - Deploying/operating services on **AWS** - Working with **AI/HPC clusters** and their compute, networking, and hardware systems - Building dashboards or applications for infrastructure operators ### Location - SF Bay Area - Toronto, Canada ### Why join Cerebras People who are serious about software make their own hardware. Cerebras is unlocking new opportunities for the AI industry with rapid growth and a breakthrough architecture. Cerebras is an equal opportunity employer committed to building an inclusive environment. **Apply:** https://www.cerebras.ai/join-us
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.