Staff Software Engineer, Quality & Reliability Platform
BuildOps · San Francisco, CA
About this role
**Staff Software Engineer, Quality & Reliability Platform** BuildOps is looking for a **Staff Software Engineer** to set and drive the company-wide technical strategy for **building, shipping, and operating reliable software**. This is a **high-impact, cross-functional** role for an engineer who treats **quality and reliability as system properties**—not a final testing phase or the responsibility of a separate team. You’ll partner across **engineering, product, infrastructure, and customer-facing organizations** to identify systemic risks, establish architectural standards, and create platform capabilities that make safe, dependable delivery the default. --- ## What You Will Do - Define and drive BuildOps’ technical strategy for **engineering quality, production reliability, and safe software delivery**. - Identify systemic sources of customer-impacting failures and lead cross-team initiatives to address **root causes**. - Establish architectural principles, engineering standards, and “**paved roads**” that make reliable design and safe delivery easier by default. - Partner with teams during system and product design to improve **resilience, operability, testability, and failure isolation** before implementation. - Build or guide shared platform capabilities for **release safety, automated validation, production feedback, test data, environment management, and developer self-service**. - Advance BuildOps’ **observability strategy** to detect regressions quickly, diagnose failures, and guide reliability investments. - Improve validation across services, data boundaries, financial workflows, and other business-critical systems. - Define meaningful measures of quality and reliability, then use them to prioritize and demonstrate improvements in customer and engineering outcomes. - Lead technical programs spanning multiple teams and organizations—aligning stakeholders and driving decisions without direct authority. - Mentor engineers and technical leaders to strengthen organizational reasoning about **risk, reliability, and quality**. - Objectively evaluate existing practices and technology, evolving or replacing them when they no longer meet needs. --- ## What Success Looks Like Success is measured by **durable improvements in outcomes**, not by the number of alerts, dashboards, tests, or frameworks created. Examples include: - Fewer customer-impacting defects and recurring classes of production failures - Greater release confidence and a lower change-failure rate - Faster detection, diagnosis, and recovery when failures occur - Shorter, more reliable feedback loops for engineers making changes - Clearer ownership and better visibility into the health of critical systems and workflows - Increased engineering velocity **without sacrificing safety or reliability** - Broad adoption of shared practices and platform capabilities without creating a centralized quality bottleneck --- ## What We Look For - Significant software engineering experience, including operating at **Staff (or equivalent)** scope on ambiguous, cross-cutting problems. - A track record of leading multi-team initiatives that improved **production reliability, delivery, platform capabilities, or engineering effectiveness**. - Strong systems thinking—connecting architecture, data integrity, operational behavior, developer workflows, and customer impact. - Experience designing and operating **distributed systems in cloud environments** (e.g., AWS). - Strong software desi
Listing freshness
CronJobs last confirmed this listing 54m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.