Senior Network & Site Reliability Engineer
Alembic · San Francisco HQ
About this role
## About Alembic Alembic is the pioneering Causal AI platform helping the world’s largest enterprises move past correlation to prove what actually drives business outcomes. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence. Alembic is backed by a **$145M Series B** from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own **NVIDIA DGX SuperPOD** built on Grace Blackwell infrastructure—one of the fastest private supercomputers in the world. ## About the Role We’re building infrastructure that must perform under real-world **scale, reliability, and security** demands. This is not a traditional “keep the lights on” role. You’ll design and operate the **global network and reliability layer** behind one of the world’s fastest private supercomputers—the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You’ll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions. ## What You’ll Do - Architect and operate **scalable, secure network architecture** for high-security requirements and large-scale ML workloads. - Own **network device configuration management** end to end, ensuring consistency and reliability across the fleet. - Improve **system and network reliability and performance** through automation, observability, and proactive capacity planning. - Implement and manage complex network protocols and connectivity, including **BGP, VPNs, and WAN circuits** and external peering. - Build and maintain comprehensive **monitoring, alerting, and incident response** (SLOs, runbooks, on-call rotations), and drive post-incident analysis and continuous improvement. - Ensure **security, compliance, and operational readiness** across network and cloud infrastructure. - Partner across engineering and data science to drive a culture of **performance and reliability**. ## What Will Help You Succeed - **8+ years** in network or infrastructure engineering, including **5+ years** in datacenter operations and/or systems and network administration. - Strong background in **network security, architecture, design, and operations**. - Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures/protocols: **BGP, QoS, MPLS, IPsec VPNs**. - Experience designing and operating modern datacenter network fabrics: **spine-leaf, EVPN/VXLAN, ECMP**. - Network automation and IaC tooling (e.g., **Ansible, Terraform, Nornir**) plus IPAM/DCIM platforms (e.g., **NetBox, Infoblox**). - **WAN engineering**: carrier circuit provisioning and external network peering. - Familiarity with **Kubernetes networking** (CNI plugins, ingress, service networking, network policy) and strong operational experience with **Linux-based production infrastructure**. - Experience with monitoring/observability stacks (e.g., **Prometheus, Grafana, Datadog, ELK, OpenTelemetry**). - Solid scripting (e.g., **Python, Bash**) to debug complex network/system issues and automate solutions; excellent cross-functional communication. ## Also Helpful - NVIDIA networking technologies: **Cumulus Linux, InfiniBand, Spectrum-X, BlueField DPUs**. - Familiarity with data-intensive platforms (**Spark
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.