Senior Site Reliability Engineer - Linux Systems & Application Observability
tastytrade · Chicago, Illinois
About this role
## Senior Site Reliability Engineer — Linux Systems & Application Observability **Company:** tastytrade (part of IG Group) **Location:** Chicago, IL — Hybrid (3 days/week in office) --- ### Role Summary Come join tastytrade as we build the reliability practice behind the brokerage platform that active options, futures, and equities traders rely on every market day. As our first **Senior Site Reliability Engineer**, you’ll harden systems behind order execution and market data delivery—designing for fault tolerance, closing telemetry gaps, and ensuring our **HashiCorp Nomad**-based service fabric scales cleanly as trading volume grows. You’ll work embedded with infrastructure and application engineering teams, contributing directly to **Ruby, Java, and Elixir** services. --- ### What You’ll Do (Job Responsibilities) - Build **self-healing, fault-tolerant infrastructure** and internal tooling to automate repetitive operational work and reduce toil. - Run **observability gap analysis** to identify blind spots in telemetry, logging, and alerting coverage—then close them so failures surface before customers feel them. - Own **scalability work** across the Nomad service fabric: capacity planning, load testing, and identifying architectural bottlenecks before incidents. - Extend the observability stack (**Prometheus, Honeycomb, OpenTelemetry**) with instrumentation needed to see failure modes. - Set **SLOs and error budgets** with multi-window burn-rate alerting for critical brokerage flows. - Mentor engineers across teams to build a lasting culture of site reliability champions. --- ### Who You Are (Skills Needed) - Hands-on experience designing **fault-tolerant, self-healing distributed systems** (not just theory—shipped in production). - Deep understanding of one or more: **distributed systems, Linux systems, cloud-native architectures, containerization**. - Experience performing **observability/telemetry gap analyses** (what’s not instrumented, not alerted on, or not visible until it’s too late). - Track record scaling systems under real production load, including **capacity planning** and **bottleneck identification**. - Hands-on experience with **OpenTelemetry, Prometheus, and Grafana**, including the ability to instrument services. - Strong **Linux internals and networking fundamentals** (TCP/IP, UDP/multicast, packet capture, flow analysis). - On-call experience on production systems and comfort building a **blameless post-incident** culture.
Listing freshness
CronJobs last confirmed this listing 23h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.