Senior Technical Duty Officer, Cloud Ops
Box · Redwood City, CA, United States
About this role
**Senior Technical Duty Officer, Cloud Ops** **About Box** Box (NYSE:BOX) is the leader in Intelligent Content Management. We help organizations fuel collaboration, manage content lifecycles, secure critical information, and transform workflows with enterprise AI. Founded in 2005, Box serves leading global organizations including JLL, Morgan Stanley, and Nationwide. **The Role** We're seeking a Senior Technical Duty Officer—a Senior Incident Commander with strong Site Reliability Engineering (SRE) expertise and proficient Python skills. You'll lead critical and blocker incidents to swift resolution while designing tools, automation, and processes for next-generation cloud operations. **Key Responsibilities** • Own live-site Critical and Blocker incidents from identification through mitigation and recovery • Triage problems, organize incident bridges, coordinate SMEs, and lead cross-functional teams • Improve incident platform tooling and automate repetitive response steps • Partner with SRE and engineering teams on Box's dependencies, failure modes, and secure workflows • Lead daily change reviews and evaluate change risk with engineering teams • Provide technical expertise in 24x7 global environments • Lead projects to enhance site resiliency, manageability, and observability • Turn incident learnings into actionable engineering improvements • Mentor and uplift the team through tabletop exercises and runbook improvements **Required Qualifications** • 5+ years in SRE, production operations, or equivalent high-scale SaaS operations • Proven Incident Commander/Technical Duty Officer experience • Strong SRE fundamentals: SLIs/SLOs, observability, blameless postmortems, toil reduction • Proficient Python for automation and tooling • Solid Linux/Unix troubleshooting and distributed systems knowledge • Networking literacy (DNS, TLS, load balancing, HTTP, routing/firewall) • Cloud environment experience (GCP preferred; AWS/Azure valued) • Container/orchestration concepts (Kubernetes or equivalent) • Excellent written and verbal communication • Proven coaching and mentoring ability **Preferred Skills** • 24x7 NOC/GTOC operations center experience • Prometheus-compatible observability, distributed tracing, synthetic monitoring, PagerDuty • Incident tooling improvement experience • Change management under pressure • Go, shell, Terraform, CI/CD expertise • Service catalog and Tier 1 journey documentation experience **Culture & Location** Minimum 3 days per week in-office. Box values diversity and encourages applications from candidates with varied backgrounds.
Listing freshness
CronJobs last confirmed this listing 8h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.