Senior Site Reliability Engineer II
Blink Health · New York, NY; Pittsburgh, PA; Seattle, WA
About this role
**Company Overview** [Blink Health](https://www.blinkhealth.com/) is the fastest growing healthcare technology company building products to make prescriptions accessible and affordable for everyone. Our products—**BlinkRx** and **Quick Save**—remove roadblocks in the prescription supply chain to improve access to critical medications and health outcomes. **BlinkRx** is the world’s first pharma-to-patient cloud with a digital concierge for patients prescribed branded medications, offering transparent low prices, free home delivery, and world-class support. --- **Responsibilities** - **Define and drive observability strategy** for IT system and process health, performance, and reliability (alerting quality, dashboards, and service health indicators). - Design and implement **software-driven solutions within the IT domain** to automate manual processes and eliminate operational complexity and toil. - Serve as a **technical leader and force multiplier**, influencing priorities and decision-making across IT (support, systems, and networks). - Own **large, ambiguous initiatives**, driving them from concept to delivery while aligning stakeholders across IT, engineering, and partner teams. - Proactively identify systemic risks and reliability gaps, **recommending and leading platform upgrades** and architectural improvements before incidents occur. - Provide **technical mentorship**, architecture guidance, and high-quality design and code reviews across IT teams. - Lead by example in **documentation and knowledge sharing** so systems and processes aren’t dependent on individual ownership. - Participate in and help mature **incident response**, escalation practices, and post-incident learning. --- **Desired Experience** - Bachelor’s or Master’s degree in Computer Science (or equivalent practical experience). - **5+ years** in site reliability engineering, infrastructure engineering, or platform engineering roles with demonstrated impact at scale. **Reliability & Troubleshooting** - Expert, methodical troubleshooting across the **entire stack** (application to kernel to network). - Strong command-line proficiency and deep expertise in **Linux systems and OS fundamentals**. - Advanced networking knowledge: **load balancing, proxies, DNS, TCP/IP, NAT**, and service-to-service communication. **Software & Automation** - Experience across multiple languages (e.g., **Python, Go, Bash**), plus familiarity troubleshooting application stacks such as **React** (strong proficiency in at least one). - Track record of **automating repetitive and complex operational work** to reduce toil and improve reliability. - Ability to design and build internal tools (Python or Go) to **standardize and scale IT practices**. - Comfortable working in an **agile environment** with disciplined testing and quality practices. **Cloud & Platform Engineering** - Cloud experience (**AWS preferred**; GCP/Azure acceptable), including managed services and production-grade architectures. - Expertise in **Kubernetes and container orchestration** (EKS, Helm), including lifecycle management and operational best practices. - Proven experience designing and implementing **observability systems** (metrics, logging, tracing, dashboards, alerting). - Deep understanding of container technologies, security scanning, secrets management, dynamic configuration, and **microservices architectures**. - Familiarity with service meshes and advanced traffic management concepts. **Infrastructur
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.