Engineering Manager, Cloud Infrastructure
Replit · Foster City, CA
About this role
## About Replit Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. ## About the Role Replit enables people to build software with AI. The infrastructure underneath that experience must make it straightforward to launch services, isolate workloads, and run reliable systems at scale. We’re hiring a **hands-on Engineering Manager** to lead **Cloud Infrastructure**—the shared infrastructure as code (IaC), networking, storage, compute, and service mesh platforms that Replit’s product and platform teams depend on. This is a **platform-building role with production accountability**. You’ll lead and grow an existing engineering team building and operating foundations including **Kubernetes, shared edge networking, service mesh, and workload identity**. ## What You’ll Do - **Own the cloud-platform roadmap**: Lead IaC, networking, storage, compute, and service mesh platforms; translate product, platform, reliability, and security needs into sequenced outcomes. - **Make infrastructure repeatable and self-service**: Build maintained IaC interfaces for services, cells, connectivity, identities, and shared resources so internal teams can provision without bespoke coordination. - **Operate what the team builds**: Own platform availability, upgrades, isolation, recovery, and incident remediation; maintain clear SLOs, sustainable on-call coverage, and primary/backup owners for critical systems. - **Stay technically engaged**: Review designs and production changes, debug difficult failure modes, and use AI coding tools (including Replit) to prototype and automate—while applying rigorous review/verification to AI-generated infrastructure changes. - **Build and grow a high-ownership engineering team**: Coach engineers, develop technical leaders, set expectations, manage performance, and hire; delegate meaningful ownership as the team grows. ## What You’ll Bring - **Demonstrated engineering management**: Led and developed engineers, made prioritization/performance decisions, hired thoughtfully, and delivered through a team. - **Software-oriented infrastructure depth**: Built and operated cloud platforms or distributed systems; can reason across infrastructure code, Kubernetes, networking, service identity, and stateful dependencies. - **Safe-change and production judgment**: Owned consequential migrations and incidents; can explain failure modes and rollback limits; knows when simplifying is better than adding another platform. - **Platform-product and engineering judgment**: Understand internal customers, create interfaces other teams adopt, and make clear tradeoffs among reliability, developer autonomy, engineering effort, and workload efficiency. ## Nice to Have - Experience with **multi-tenant, cellular, regional, or dedicated enterprise infrastructure**. - Familiarity with **GCP/GKE**, **Terraform** (or similar IaC), **Cloudflare**, **Envoy/Istio**, **SPIFFE/SPIRE**, and managed data services. - Experience with **large-fleet rightsizing**, infrastructure consolidation, or migrating CI compute without disrupting developer workflows. - Track record using **AI tools** to increase output while preserving production safeguards. ## Benefits (Full-Time) - 💰 Competitive Salary & Equity - 💹 401(k) Program with a 4% match (US Only) - ⚕️ Health, Dental, Vis
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.