CronJobs

devops-sre jobs

Senior DevOps Engineer

Alpaca · Remote - Americas

remoteseniorPosted Aug 28, 2026GCPTerraformKubernetesGKEHelmDockerPrometheusGrafana

Apply on the employer site

About this role

**Who We Are** Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. Amongst our subsidiaries, Alpaca is a licensed financial services company serving hundreds of financial institutions across 40 countries with institutional-grade APIs—broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges—totalling over 10 million brokerage accounts. We’re a diverse global team of experienced engineers, traders, and brokerage professionals working toward our mission of opening financial services to everyone on the planet. We’re deeply committed to open-source contributions and building a vibrant developer community. **Role: Senior DevOps Engineer** Design, build, and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You’ll have autonomy to implement solutions against clearly defined goals—and a real voice in shaping those goals with the team. We’re not hiring a specialist in a single tool. We’re looking for a well-rounded infrastructure engineer who thinks in cloud architecture and Infrastructure-as-Code, with a Platform-as-a-Product mindset—measuring success by how quickly and safely the rest of engineering can ship, and treating manual toil as a bug to be engineered away. You’ll operate data stores (PostgreSQL, message brokers) at an operator level and partner with SRE and database specialists. **Things You Get To Do** - **Design and evolve** our cloud architecture on **GCP** (networking, interconnects, IAM, high-availability topology) and express it entirely as code with **Terraform**, following **GitOps** as a first principle. - **Build and own** CI/CD pipelines for IaC (plan/review/test/apply) with **Policy-as-Code** guardrails, drift detection, and progressive rollout. - **Advance Platform-as-a-Product**: build self-serve capabilities and paved paths so engineers can provision what they need via a golden path. - **Strengthen observability** across **Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager** (metrics, logs, traces, alerting) so the platform is easy to run and reason about. - **Operate GKE clusters** and infrastructure services running on them (Helm-packaged workloads, message brokers like **RabbitMQ** and **IBM MQ**, and data stores). - Participate in our **Follow-The-Sun on-call** model: triage alerts, join/declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and post-actions. - **Embed SRE practices** (SLIs/SLOs, error budgets, capacity planning) into how Core Infrastructure builds and operates. **Who You Are (Must-Haves)** - **5+ years** in DevOps, Platform/Infrastructure, or SRE with a track record operating large-scale, high-availability, high-performance production systems. - Deep hands-on experience designing cloud architecture on **GCP** (landing zones, networking, IAM, high-availability topology). - Strong **Infrastructure-as-Code** skills with **Terraform**, structuring large codebases across multiple environments; **GitOps** as a first principle and **least-privilege** by default. - Proven experience building **CI/CD pipelines for IaC** (automated plan/apply, code review, Policy-as-Code, drift detection, safe rollout). - Significant production experience with **Kubernetes** (ideally **GKE**) and deploying workloads with **Helm**. - Solid L3/L4–L7 networking fundamentals (VPCs,

Listing freshness

CronJobs last confirmed this listing 8h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord