CronJobs

backend jobs

AI Infrastructure Engineer, pAGI

OpenAI · San Francisco

hybridunknown$266,000–$500,000Posted Sep 14, 2026distributed systemsinferenceGPU performancecompute schedulingobservabilityhealth monitoring

Apply on the employer site

About this role

## About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Work spans: - Distributed training infrastructure - Inference and grading platforms - Compute scheduling - Research tooling You’ll partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. ## About the Role We’re looking for an **AI Systems Engineer** to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. ## What You’ll Do - Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. - Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. - Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. - Diagnose bottlenecks across training, inference, and orchestration, working across teams to improve end-to-end performance. - Build self-service tools, automated validation, and observability to help researchers launch experiments, diagnose issues, and compare results with less manual intervention. ## You Might Thrive If You - Are excited about the potential of personal AGI and want to build the infrastructure that enables it. - Have strong software engineering fundamentals and experience building or operating large-scale distributed systems. - Have experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling. - Are highly self-motivated and comfortable taking ownership of open-ended problems. - Enjoy debugging across system boundaries and using measurements to guide improvements in performance and reliability. ## About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of AI capabilities and seek to safely deploy them through our products. OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities. (For full policy details and background check information, please refer to the original posting.)

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord