CronJobs

backend jobs

AI Fleet Platform Software Engineer

Cerebras · Sunnyvale, CA

hybridseniorPosted Oct 2, 2026GoPythonSQLProtocol BuffersKubernetesLinuxAWSgRPC

Apply on the employer site

About this role

## AI Fleet Platform Software Engineer ### About the role Cerebras Systems builds large-scale AI compute infrastructure and operates AI clusters across multiple data centers. As the fleet grows, we need a senior software engineer to build the platforms and tools that help teams monitor fleet health, respond to issues, and keep compute capacity available. You’ll own substantial engineering work across **backend services**, **operational tools**, and **user-facing applications**, partnering with Cluster Operations, infrastructure, inference, security, and data center teams to turn real operational challenges into reliable software at scale. ### Responsibilities - Build and operate software for managing large fleets of AI clusters - Provide operators clear, actionable views of cluster health, capacity, performance, and ongoing issues - Develop services and integrations that bring together data and workflows from multiple infrastructure systems - Automate repetitive work and improve tools for investigating incidents and restoring service - Design systems that remain reliable as the fleet grows, including through component and site failures - Work closely with platform users to understand needs and make practical product/engineering decisions - Lead projects from initial design through production, measure impact, and use operational feedback to improve ### Skills & Requirements - **12+ years** building and operating production distributed systems or large-scale infrastructure software - Strong **Go** skills (regular work with Go, Protocol Buffers, and SQL). **Python** is a plus - Experience building **control planes**, **fleet management systems**, or **operational platforms** - Experience designing **APIs and data models** for multiple consumers, with thoughtful versioning/compatibility - Experience with **asynchronous systems** supporting idempotency, concurrency, and reconciliation while preventing state drift - Strong testing practices, especially for software that can change the state of running production clusters - Experience with **Linux**, **containers**, and **Kubernetes** - Experience with durable workflow execution, event streaming, or high-cardinality time-series telemetry - Sound judgment around **reliability, security, and observability** (including dependency unavailability) - Ability to take an ambiguous problem from design to production and improve using operational feedback ### Preferred Experience - Infrastructure software for thousands of hosts across multiple sites - Incident response, hardware health, or capacity management software - Deploying/operating services on **AWS** - Working with **AI/HPC clusters** and their compute, networking, and hardware systems - Building dashboards or applications for infrastructure operators ### Location - SF Bay Area - Toronto, Canada ### Why join Cerebras People who are serious about software make their own hardware. Cerebras is unlocking new opportunities for the AI industry with rapid growth and a breakthrough architecture. Cerebras is an equal opportunity employer committed to building an inclusive environment. **Apply:** https://www.cerebras.ai/join-us

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord