CronJobs

ai-ml jobs

Senior Machine Learning Engineer

Air · Arlington, Virginia, United States; Pittsburgh, Pennsylvania, United States; Remote

remoteseniorPosted Sep 9, 2026PythonKubernetesAWSGCPAzureLLMOpsLLM fine-tuningModel serving

Apply on the employer site

About this role

**Senior Machine Learning Engineer** **Company Overview** Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. We build an AI-native platform—**Air Enterprise Readiness**—that aligns development, production, delivery, and sustainment into one coordinated execution system for government agencies and industrial suppliers. **Role Summary** Air is seeking an experienced **Senior Machine Learning Engineer** to join our AI/ML team and build the infrastructure that powers the development, evaluation, deployment, and continuous improvement of our language models and AI systems. As AI capabilities expand, this role will own critical parts of the lifecycle—especially **LLMOps**, **fine-tuning infrastructure**, **model evaluation**, **dataset pipelines**, **experiment management**, **model serving**, and **production observability**. **What You’ll Do** - Design and build **LLMOps infrastructure** for production language models. - Build scalable **training and fine-tuning infrastructure** for commercial and open-weight models. - Develop pipelines for **supervised fine-tuning**, **parameter-efficient fine-tuning** (e.g., LoRA/QLoRA), **preference optimization**, and other post-training techniques. - Build infrastructure for **distributed training** and **GPU-accelerated** ML workloads. - Develop data pipelines for **training, fine-tuning, evaluation, and synthetic data generation**. - Create systems for **dataset versioning, lineage, quality validation, transformation, and reproducible experimentation**. - Build **experiment management** infrastructure to compare models, datasets, hyperparameters, prompts, and training techniques. - Implement automated **evaluation pipelines** to determine production readiness. - Design **model registries**, artifact management, versioning, and promotion workflows. - Build and operate scalable **model-serving/inference** infrastructure for open-weight and fine-tuned models. - Develop abstractions so product/AI teams can use multiple models and inference providers without tight coupling to a single vendor. - Build **observability** for training and inference (metrics, tracing, logging, utilization, quality, latency, throughput, and cost). - Optimize workloads for **GPU utilization, throughput, latency, reliability, and infrastructure cost**. - Build automated workflows for **deployment, rollback, canarying, and production validation**. - Investigate failures across data pipelines, training jobs, inference services, distributed systems, and production environments. - Evaluate emerging models, training techniques, inference frameworks, and ML infrastructure to improve production systems. - Partner with teams building **agentic systems** to provide model, evaluation, and training infrastructure for continuous improvement. **Location & Travel** - Full-time role based in **Pittsburgh, PA**, or **Remote** opportunities. - May require up to **25% travel**, including periodic trips to **Pittsburgh, PA** and **Arlington, VA**. **Qualifications** - **U.S. Citizenship is required** **Required Skills** - 5+ years building **production ML systems**, **ML infrastructure**, **distributed systems**, or similar technical systems. - Deep experience designing/building/operating **production ML infrastructure or ML platforms**. - Experience with the LLM lifecycle: **data preparation, training, evaluation, deployment, inference, monitoring, and ite

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord