Senior Software Engineer, Platform & Infrastructure - Riot Technology
Riot Games · Los Angeles, USA; Mercer Island, USA; Redwood City, USA
About this role
**Platform & Infrastructure — Senior Software Engineer (Riot Technology)** Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. You’ll partner across disciplines to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players. As a **Senior Platform & Infrastructure Engineer** on the **Riot Technology** team, you will design, build, and operate core infrastructure and ML platforms—especially the computing and orchestration platforms that power large-scale distributed training of agents (e.g., RL, IL), simulation environments, and policy evaluation. You’ll also drive production-grade CI/CD, infrastructure-as-code, observability, and developer tooling, closing critical infrastructure gaps and improving standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning. **Responsibilities** - Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation. - Design infrastructure for running simulation environments at scale (parallel rollouts, data collection, training, and evaluation). - Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments. - Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity. - Build observability (monitoring, alerting, health indicators) and SLO-aligned dashboards for infrastructure and ML workloads. - Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows. - Support MLOps workflows (automated training pipelines, model artifact management, experiment tracking, reproducible ML lifecycle operations). - Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles. **Required Qualifications** - Bachelor’s degree in Computer Science (or related) or equivalent practical experience. - 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles. - Experience operating distributed systems in production and keeping them healthy under real load. - Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling. - Experience with GPU compute infrastructure (scheduling, multi-node orchestration, resource optimization for long-running training workloads). - Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals. **Desired Qualifications** - Familiarity with MLOps workflows (model versioning, pipeline orchestration, experiment tracking, artifact management, reproducible ML workflows). - Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools. - Passion for games and player experience. **Perks (Highlights)** - Work/life balance with open paid time off and flexible schedules. - Medical, dental, and life insurance; parental l
Listing freshness
CronJobs last confirmed this listing 2d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.