CronJobs

devops-sre jobs

Site Reliability Engineer

Blaxel · San Francisco

onsitemid$175,000–$250,000Posted Mar 3, 2026PythonGoRustAWSGCPKubernetesTerraformPrometheus

Apply on the employer site

About this role

**Site Reliability Engineer** **About the Role** We’re looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You’ll build and operate the core systems that power agentic AI at scale—keeping our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests. This role is highly technical and execution-heavy. You’ll own reliability end-to-end, including observability, performance tuning, incident operations, infrastructure health, and automation. **What You’ll Do** - Architect, operate, and continuously improve the core infrastructure powering our **25ms cold-start compute engine** - Build and evolve our **observability stack** (metrics, traces, logs) to detect issues before users do - Define, monitor, and drive **SLOs/SLIs** across key system surfaces - Lead incident response with rigor: **root cause analysis** and **post-mortems**, driving systemic fixes - Design and implement **self-healing, automated operational systems** to eliminate toil - Work across **compute, networking, storage, and sandboxed execution** layers to tune performance under extreme workloads - Build automation and tooling (often with **AI agents**) for operations, debugging, capacity planning, and failure prediction - Stress-test systems using **load testing, chaos engineering, and performance benchmarking** - Own infrastructure-layer **security best practices** (sandboxing, network isolation) - Partner with platform engineers to ensure reliability is built into new features from day one **Who You Are** - Deeply technical across systems, cloud, networking, and distributed computing; you debug real failures - AI-fluent operator: you understand how AI systems behave under scale and their resource patterns - Builder at heart: you want to invent new reliability systems in a zero-to-one environment - High-velocity execution with strong judgment and a track record of shipping reliable systems quickly - Automation-first mindset (you hate repeated manual work) - Calm under pressure during incidents—clear, precise, and ownership-driven - Data-driven: you measure latency, tail behavior, resource efficiency, and reliability trends **Required Skills** - **3+ years** in SRE, DevOps, or infrastructure engineering - Proficiency in at least one language: **Go, Rust, or Python** - Hands-on experience with a major cloud provider (**AWS, GCP**) - Solid knowledge of **Linux**, networking fundamentals, and distributed systems - Experience with **bare-metal servers and datacenter operations**, including provisioning and high-throughput networking - Experience with **Kubernetes** (or similar orchestrators) - Familiarity with observability stacks: **Prometheus, Grafana, ELK, Datadog** - Experience building/maintaining **CI/CD pipelines** (GitHub Actions, GitLab CI, Jenkins) - Strong debugging, problem-solving, and incident-management skills **Preferred** - Infrastructure-as-code: **Terraform** or **Pulumi** - Service mesh or **API gateway** technologies - Chaos engineering / resiliency-testing frameworks - Security best practices for cloud environments - Experience in high-growth or high-availability environments **Bonus** - Serverless compute systems - Sandboxed execution environments - Ultra-low-latency runtime engineering - Distributed key-value stores and databases - Chaos engineering - Systems-level programming (Rust/Go) - Deep generative AI infrast

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord