Senior Production Operations Engineer
Sony Interactive Entertainment Global · United States, San Diego, CA
About this role
## Senior Production Operations Engineer **Why Sony Interactive Entertainment?** Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. As the company behind the PlayStation brand, SIE delivers cutting-edge hardware and network services to more than 100 million people, and is home to beloved intellectual properties (IP) worldwide. As part of the **Production Operations Engineering** team within the **Platform Technology** group, you’ll help keep key user experiences **available, resilient, and high-performing** across time zones and critical business periods. The team responds when production services need support through **rotational on-call, incident response, and urgent operational escalations**. You’ll also lead technical initiatives that improve **process, technology, and production operations** for millions of users. This is a **senior individual contributor (P4)** role: highly independent, hands-on, influential across partner teams, and accountable for meaningful production operations capabilities. The ideal candidate brings deep production experience, strong automation instincts, modern cloud/container operations expertise, and curiosity about using **AI-assisted engineering workflows** to improve operational quality and delivery speed. ### Responsibilities - Own application operations and production support for internal and public-facing services in **AWS** and **container** environments, focusing on **availability, resiliency, scalability, performance, and security**. - Drive production readiness for new services and features, including **provisioning, automation, monitoring, alerting, dashboards, runbooks, rollback planning, and operational acceptance criteria**. - Build scripts, tools, and repeatable workflows to reduce toil, improve incident response, and standardize operational practices. - Improve **observability** and **event correlation** using actionable service health signals, dashboards, alerts, and incident workflows to reduce **MTTD/MTTR**. - Support and improve incident management practices using tools such as **Slack, JIRA, ServiceNow, BigPanda**, and related escalation/notification systems. - Partner with **SRE, data services, CI/CD, service engineering, platform hosting, and product teams** to improve end-to-end reliability and operational readiness. - Drive **performance, capacity, and cost optimization** using **AWS, Kubernetes, EKS/UKS, autoscaling**, and related cloud patterns. - Use **AI-assisted/agentic engineering tools** (e.g., Codex, Claude, Cursor, or similar) to improve operational quality and delivery speed.
Listing freshness
CronJobs last confirmed this listing 3h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.