Senior Infrastructure Engineer, Storage Platform
Cloudflare · In-Office
About this role
## Senior Infrastructure Engineer, Storage Platform **About Us** At Cloudflare, we’re on a mission to help build a better Internet. We protect and accelerate Internet applications online—without adding hardware, installing software, or changing a line of code. **Available Locations** Austin (US) • Seattle (US) • London (UK) --- ## About the Role Emerging Technologies & Incubation (ETI) builds and launches new products on Cloudflare’s global network. Within ETI, you’ll join the **Storage Infrastructure** team to build and operate a shared storage platform for stateful products such as **R2**, **Workers KV**, and **Durable Objects**. You’ll manage underlying storage hardware, distributed databases, and object storage clusters—covering fleet lifecycle automation, capacity, hardware validation, failure-domain planning, observability, and production operations. In this role, you’ll build platform software and automation to make fleet operations **safe, repeatable, and scalable**. --- ## Responsibilities - Design and build automation and operator tooling for provisioning, configuring, expanding, upgrading, and decommissioning storage hardware, distributed databases, and object storage clusters. - Engineer globally distributed storage fleets to tolerate hardware failures, network disruption, capacity pressure, and recovery events across multiple failure domains. - Develop observability, alerting, and safety controls to help engineers understand fleet health and make auditable production changes. - Diagnose incidents across Linux, storage, networking, and distributed systems; participate in on-call and implement durable fixes that reduce operational toil. - Use AI tools throughout development and operations to accelerate debugging, incident triage, root-cause analysis, and toil reduction—while verifying outputs before taking production action. - Validate new storage hardware and characterize performance during normal operation, failures, rebuilds, and other recovery conditions. - Partner with R2, Workers KV, Durable Objects, Network, Capacity Planning, Performance Engineering, and Infrastructure Operations to turn service requirements into infrastructure capabilities. --- ## Desirable Skills, Knowledge, and Experience - Experience designing, building, and operating infrastructure platforms or large-scale production fleets, including automation of lifecycle operations. - Strong Linux systems knowledge; experience troubleshooting across compute, storage, and networking layers. - Experience operating distributed systems in production, including observability, incident response, reliability improvements, and safe change management. - Ability to write maintainable software in at least one language such as **Go**, **Rust**, or **Python**. - Strong written and verbal communication skills; experience collaborating across engineering teams and technical stakeholders. --- ## Bonus Points - Experience operating distributed databases or object storage systems; knowledge of storage fundamentals (filesystems, SSD behavior, replication, rebuild dynamics). - Experience with bare-metal provisioning, hardware lifecycle management, or storage hardware validation. - Experience with infrastructure-as-code, configuration management, or observability systems (e.g., Terraform, SaltStack, Prometheus, Grafana). - Experience with capacity modeling, workload placement, failure-domain planning, or cross-datacenter networking. --- ## Compensation & Equity - **C
Listing freshness
CronJobs last confirmed this listing 15h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.