Senior DevOps Engineer
Alpaca · Remote - Americas
About this role
**Who We Are** Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. Amongst our subsidiaries, Alpaca is a licensed financial services company serving hundreds of financial institutions across 40 countries with institutional-grade APIs—broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges—totalling over 10 million brokerage accounts. We’re a diverse global team of experienced engineers, traders, and brokerage professionals working toward our mission of opening financial services to everyone on the planet. We’re deeply committed to open-source contributions and building a vibrant developer community. **Role: Senior DevOps Engineer** Design, build, and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You’ll have autonomy to implement solutions against clearly defined goals—and a real voice in shaping those goals with the team. We’re not hiring a specialist in a single tool. We’re looking for a well-rounded infrastructure engineer who thinks in cloud architecture and Infrastructure-as-Code, with a Platform-as-a-Product mindset—measuring success by how quickly and safely the rest of engineering can ship, and treating manual toil as a bug to be engineered away. You’ll operate data stores (PostgreSQL, message brokers) at an operator level and partner with SRE and database specialists. **Things You Get To Do** - **Design and evolve** our cloud architecture on **GCP** (networking, interconnects, IAM, high-availability topology) and express it entirely as code with **Terraform**, following **GitOps** as a first principle. - **Build and own** CI/CD pipelines for IaC (plan/review/test/apply) with **Policy-as-Code** guardrails, drift detection, and progressive rollout. - **Advance Platform-as-a-Product**: build self-serve capabilities and paved paths so engineers can provision what they need via a golden path. - **Strengthen observability** across **Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager** (metrics, logs, traces, alerting) so the platform is easy to run and reason about. - **Operate GKE clusters** and infrastructure services running on them (Helm-packaged workloads, message brokers like **RabbitMQ** and **IBM MQ**, and data stores). - Participate in our **Follow-The-Sun on-call** model: triage alerts, join/declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and post-actions. - **Embed SRE practices** (SLIs/SLOs, error budgets, capacity planning) into how Core Infrastructure builds and operates. **Who You Are (Must-Haves)** - **5+ years** in DevOps, Platform/Infrastructure, or SRE with a track record operating large-scale, high-availability, high-performance production systems. - Deep hands-on experience designing cloud architecture on **GCP** (landing zones, networking, IAM, high-availability topology). - Strong **Infrastructure-as-Code** skills with **Terraform**, structuring large codebases across multiple environments; **GitOps** as a first principle and **least-privilege** by default. - Proven experience building **CI/CD pipelines for IaC** (automated plan/apply, code review, Policy-as-Code, drift detection, safe rollout). - Significant production experience with **Kubernetes** (ideally **GKE**) and deploying workloads with **Helm**. - Solid L3/L4–L7 networking fundamentals (VPCs,
Listing freshness
CronJobs last confirmed this listing 8h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.