CronJobs

devops-sre jobs

Network Operations Engineer, AI Networking

OpenAI · San Francisco

remoteunknown$157,000–$302,000Posted Sep 15, 2026PythonTerraformAWSAzureCisco NX-OSArista EOSJuniper JunOSPrometheus

Apply on the employer site

About this role

## About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute’s data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. ## About the Role We’re seeking an **Infrastructure Operations Engineer** to operate and improve large-scale Ethernet fabrics that support **GPU clusters, storage systems, and management infrastructure**. This role combines hands-on production operations with **automation, observability, and incident response** across a global AI network. You’ll move comfortably from **physical-layer troubleshooting** to **routing and fabric behavior**, execute changes, and perform **root-cause analysis**. You’ll partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. ## Key Responsibilities - Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute’s data centers. - Monitor, troubleshoot, and resolve network incidents while meeting **SLOs**, reducing **MTTD**, and minimizing **MTTR**. - Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. - Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. - Manage the hardware lifecycle (switch/optics replacements, RMA coordination, software upgrades, preventive maintenance). - Support new AI cluster deployments, data center expansions, and infrastructure migrations with deployment and engineering teams. - Partner with CSPs, colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure. - Perform **root-cause analysis (RCA)** and drive permanent corrective actions to eliminate recurring issues. - Build and maintain monitoring, telemetry, dashboards, and alerting to improve observability and proactive detection. - Develop and improve operational runbooks, playbooks, troubleshooting documentation, and SOPs. - Automate repetitive operational tasks using **Python** and infrastructure automation frameworks to reduce toil. - Continuously identify opportunities to improve reliability, scalability, operational maturity, and engineering efficiency. ## Qualifications - Bachelor’s degree in Computer Science, Network Engineering, or related field, **or equivalent practical experience**. - **5+ years** operating large-scale data center, cloud, AI, or HPC network infrastructure. - Experience supporting production network environments with **high-availability** requirements. - Hands-on experience with one or more: **Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, Juniper JunOS**. - Strong knowledge of **Layer 2/Layer 3 networking** and protocols/technologies including **BGP, OSPF, ECMP, MLAG, LACP, VRFs, VLANs**. - Experience troubleshooting physical infrastructure: **fiber optics, transceivers, DAC/AOC cables, high-speed Ethernet links**. - Experience with software upgrades, hardware maintenance, and production change management. - Excellent analytical/troubl

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord