Site Reliability Engineer, Platform Infrastructure (Foundations)
Anyscale · San Francisco
About this role
**Site Reliability Engineer, Platform Infrastructure (Foundations)** **About Anyscale:** Anyscale is democratizing distributed computing by commercializing Ray, an open-source project powering scalable ML at companies like OpenAI, Uber, Spotify, and Instacart. Backed by Andreessen Horowitz, NEA, and Addition with $250M+ raised. **About the Role:** Join the Infrastructure team to build the scalable, secure backbone enabling distributed AI applications. You'll work on both control plane (cluster orchestration, scheduling, access) and data plane (high-performance workload execution), designing and optimizing critical infrastructure for Anyscale's cloud platform. **Key Projects:** • Design and scale services orchestrating Ray clusters across cloud/on-prem (VM and Kubernetes) • Optimize control plane for large-scale distributed AI/ML workloads • Build intelligent scheduling and resource management for heterogeneous compute • Enhance reliability, performance, scalability, and observability • Support accelerator integration (GPUs, TPUs) • Manage container images and dependency resolution • Participate in design discussions and on-call support **Requirements:** • Bachelor's in Computer Science/Engineering or equivalent experience • 3+ years writing production code • Experience building highly available, scalable distributed systems • Cloud-native expertise (AWS, Azure, GCP) and Kubernetes • Deep understanding of networking, security, and authentication • Familiarity with observability stacks (Prometheus, Grafana) • Proficiency in Go and Python • Knowledge of Linux kernel, file systems, and containers **Equal Opportunity Employer** | E-Verify Company
Listing freshness
CronJobs last confirmed this listing 23h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.