CronJobs

backend jobs

Software Engineer, Inference Platform

Cerebras · Sunnyvale, CA

onsitemidPosted Jun 20, 2026GoC++KubernetesTLSmTLSCI/CD

Apply on the employer site

About this role

## Software Engineer, Inference Platform ### About Cerebras Cerebras Systems builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (over 10x faster than GPU-based hyperscale cloud inference). This performance shift is transforming AI application experiences by unlocking real-time iteration and additional agentic computation. Cerebras works with leading model labs, global enterprises, and AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras to deploy 750MW of scale for ultra high-speed inference: https://openai.com/index/cerebras-partnership/ ### About the Role Join the **Inference Platform** team. You’ll help build the orchestration layer that runs inference on Cerebras datacenter clusters—connecting cloud components with machine learning services. You’ll often be the first team to tackle problems that haven’t been solved yet, driving solutions across **Kubernetes operators**, **service security policies**, and **CI/CD**. If you’re interested in building the next-generation architecture of a globally distributed inference platform, we’d like to talk. ### Responsibilities - Design, develop, test, and maintain production software across testing, continuous development, observability, security, networking, debugging, and productionization. - **Platform Direction:** Shape technical direction for the Inference Platform, including Kubernetes custom resource definitions, failure domains, service boundaries, and system evolution; own roadmaps for major technical areas. - **Reliability & Performance:** Architect active-active systems with rapid failover, graceful degradation, and clear SLOs. Improve latency, throughput, capacity efficiency, and resilience under unpredictable demand. - **Execution on Critical Paths:** Write and review production code in the most important parts of the platform. Make high-consequence architectural decisions and set the technical bar through design and code reviews. - **Production Leadership:** Lead hardest production issues and cross-system bottlenecks. Drive observability, incident response, capacity planning, and post-incident improvements with operational rigor. - **Technical Influence:** Partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs and align on shared technical decisions. ### Skills & Qualifications - 3+ years of software engineering experience; experience building and operating large-scale distributed systems or cloud infrastructure. - Experience with distributed systems (ideally Kubernetes). - Experience building highly available, latency-sensitive systems at scale. - Security experience (certificates, TLS, mTLS). - Experience optimizing latency, throughput, and efficiency in high-QPS systems (TTFT and tail-latency reduction is a strong plus). - Strong proficiency in backend or systems languages such as **Go** or **C++**. ### Preferred Qualifications - Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads. ### Location Open to **Sunnyvale** or **Toronto**. ### Why Join Cerebras People who are serious about software make their own hardware. With dozens of model releases and rapid growth, Cerebras is at an inflection point—backed by a breakthrough architecture and a culture that values learning, growth, and support. **Five reasons team members joined:** 1. Build a breakthrough AI platform beyond GPU co

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord