CronJobs

devops-sre jobs

Senior Site Reliability Engineer - Linux Systems & Application Observability

tastytrade · Chicago, Illinois

hybridsenior$180,000–$200,000Posted Sep 4, 2026LinuxJavaRubyElixirHashiCorp NomadOpenTelemetryPrometheusGrafana

Apply on the employer site

About this role

## Senior Site Reliability Engineer — Linux Systems & Application Observability **Company:** tastytrade (part of IG Group) **Location:** Chicago, IL — Hybrid (3 days/week in office) --- ### Role Summary Come join tastytrade as we build the reliability practice behind the brokerage platform that active options, futures, and equities traders rely on every market day. As our first **Senior Site Reliability Engineer**, you’ll harden systems behind order execution and market data delivery—designing for fault tolerance, closing telemetry gaps, and ensuring our **HashiCorp Nomad**-based service fabric scales cleanly as trading volume grows. You’ll work embedded with infrastructure and application engineering teams, contributing directly to **Ruby, Java, and Elixir** services. --- ### What You’ll Do (Job Responsibilities) - Build **self-healing, fault-tolerant infrastructure** and internal tooling to automate repetitive operational work and reduce toil. - Run **observability gap analysis** to identify blind spots in telemetry, logging, and alerting coverage—then close them so failures surface before customers feel them. - Own **scalability work** across the Nomad service fabric: capacity planning, load testing, and identifying architectural bottlenecks before incidents. - Extend the observability stack (**Prometheus, Honeycomb, OpenTelemetry**) with instrumentation needed to see failure modes. - Set **SLOs and error budgets** with multi-window burn-rate alerting for critical brokerage flows. - Mentor engineers across teams to build a lasting culture of site reliability champions. --- ### Who You Are (Skills Needed) - Hands-on experience designing **fault-tolerant, self-healing distributed systems** (not just theory—shipped in production). - Deep understanding of one or more: **distributed systems, Linux systems, cloud-native architectures, containerization**. - Experience performing **observability/telemetry gap analyses** (what’s not instrumented, not alerted on, or not visible until it’s too late). - Track record scaling systems under real production load, including **capacity planning** and **bottleneck identification**. - Hands-on experience with **OpenTelemetry, Prometheus, and Grafana**, including the ability to instrument services. - Strong **Linux internals and networking fundamentals** (TCP/IP, UDP/multicast, packet capture, flow analysis). - On-call experience on production systems and comfort building a **blameless post-incident** culture.

Listing freshness

CronJobs last confirmed this listing 23h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord