Lead Support Engineer
Wppmedia · Los Angeles, United States; San Francisco, United States
About this role
**About WPP Media** WPP is the trusted growth partner for the world’s leading brands. With exceptional talent, trusted data and intelligence, and world-class partnerships—united by our pioneering agentic marketing platform, **WPP Open**—we help clients navigate change, capture opportunity, and deliver transformational growth. **WPP Media** is WPP’s AI-driven media operating unit, bringing together media, data, and partnerships to deliver creative personalisation at scale. Connected through WPP Open and powered by Open Intelligence, clients see exactly where, how, and why their media investment is working. For more information: https://wppmedia.com --- **Role Summary and Impact** As **Lead Support Engineer** within WPP Media, you will own **production stability, observability, and system health** for a mission-critical global **campaign governance and compliance** platform running natively on **Google Cloud Platform (GCP)**. You’ll provide advanced **L2+ and L3 support**, investigate complex production issues, implement minor **Python** fixes, lead **root cause analysis**, and improve the reliability of distributed systems. You’ll work closely with the **EMEA Engineering Tech Lead**, **Product Owner**, and **Quality Assurance** partners to connect production insights to technical roadmap priorities. This role combines hands-on **Site Reliability Engineering (SRE)** with technical squad leadership—building on an existing monitoring/alerting foundation to strengthen observability, automate runbooks, reduce manual operational effort, and help shape the future support squad for a dedicated product used across global advertising campaigns. --- **Key Responsibilities** - Own **production stability**, **system health**, and **observability** for the platform running on **GCP**. - Diagnose and resolve complex, intermittent, high-priority incidents across **application, database, infrastructure, and networking** layers. - Read, debug, and implement minor fixes and patches in the existing **Python** production codebase. - Define and advance the **observability strategy** using **GCP Cloud Logging**, **Cloud Monitoring**, **Error Reporting**, **Prometheus/PromQL**, and appropriate **SLIs/SLOs**. - Lead **incident response**, **post-incident reviews**, and end-to-end **root cause analysis**, partnering with the EMEA Engineering team on permanent remediation. - Monitor execution flows, performance, data refreshes, automated checks, and recurring compliance reporting. - Identify opportunities to automate manual support activity; develop and maintain **runbooks** to reduce operational toil. - Provide technical leadership to the **L2+ support squad** through coaching, knowledge sharing, prioritization, and effective handovers. - Partner with **Product, Engineering, and QA** to assess operational risk, improve release/change management, and influence technical roadmap decisions. --- **Skills and Experience** - Advanced education in computer science/software engineering/IT (or equivalent practical experience). - Strong, current **Python** proficiency (read/debug/implement fixes in a shared production codebase). - Deep hands-on experience supporting enterprise systems on **GCP**, including: - Compute Engine, Google Kubernetes Engine, Cloud SQL, BigQuery, Pub/Sub, Firestore, Cloud Functions - Identity and Access Management, VPC networking - Experience with complex production troubleshooting and **root cause analysis** in distributed systems, i
Listing freshness
CronJobs last confirmed this listing 11h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.