Staff/Senior Software Engineer - Data Observability
Acryl Data · Palo Alto, California, United States
About this role
**Data Observability — Staff/Senior Software Engineer (SRE Tech Lead)** **About DataHub** DataHub is an AI & Data Context Platform adopted by 3,000+ enterprises, including Apple, CVS Health, Netflix, and Visa. Built with a thriving open-source community (13,000+ members), DataHub’s metadata graph provides deep context for AI and data assets with best-in-class scalability and extensibility. **About the Role** We’re seeking an experienced **Site Reliability Engineering (SRE) Tech Lead** to drive the reliability, scalability, and operational excellence of our platform offerings. You’ll lead technical initiatives across **DataHub Cloud** and an emerging **enterprise deployment** solution that gives customers more control and flexibility to run DataHub in their preferred environments. --- ## Key Responsibilities ### Technical Leadership & Architecture - Design and implement robust, scalable infrastructure solutions for DataHub Cloud and enterprise deployments - Lead the technical vision for multi-cloud deployment strategies and distributed system integrations - Architect monitoring, observability, and alerting systems across diverse environments - Drive best practices for infrastructure as code, configuration management, and deployment automation ### Enterprise Platform Development - Partner with product and engineering teams to influence advanced deployment capabilities - Collaborate cross-functionally to build seamless installation, upgrade, and rollback processes across environments - Influence and help implement comprehensive monitoring and health check systems for distributed deployments - Help develop self-healing and automated remediation capabilities ### Platform Reliability & Operations - Establish and maintain **SLAs/SLOs** for cloud and enterprise offerings - Lead incident response and post-mortem processes to drive continuous improvement - Implement chaos engineering practices to proactively identify weaknesses - Optimize performance, capacity planning, and cost efficiency ### Team Leadership & Collaboration - Mentor and guide SRE engineers; collaborate with platform engineering teams - Work with product, engineering, and customer success to ensure reliable delivery - Improve on-call practices, runbooks, and knowledge sharing - Drive cross-functional initiatives to improve overall system reliability --- ## Required Qualifications - **8+ years** in SRE, Platform Engineering, or DevOps - **3+ years** technical leadership experience managing engineering teams - Strong expertise with **AWS, GCP, Azure** and infrastructure automation - Proficiency with **Docker** and **Kubernetes** (containerization + orchestration) - Experience with **Terraform, CloudFormation, or Pulumi** - Strong programming skills in **Python, Java, or similar** - Deep understanding of monitoring/observability tools (e.g., **Prometheus, Grafana, Datadog**) - Experience with **CI/CD** and deployment automation - Strong knowledge of networking, security, and database operations in cloud environments --- ## Preferred Qualifications - Experience building/operating **multi-tenant SaaS** platforms - Background building customer-facing deployment/management tools - Knowledge of data infrastructure and metadata management systems - Experience with **service mesh** and **microservices** architectures - Customer-facing technical experience with enterprise clients - Experience with data governance or data catalog platforms --- ## What You’ll Build - A robust ma
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.