CronJobs

ai-ml jobs

Senior Research Engineer, LLM Training & Post-Training

Lightning AI · New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States

remotesenior$165,000–$165,000Posted Aug 12, 2026PyTorchPythonTransformersCUDADistributed SystemsFSDPDeepSpeedvLLM

Apply on the employer site

About this role

## Senior Research Engineer, LLM Training & Post-Training **About Lightning AI** Lightning AI is the company behind PyTorch Lightning. We build an end-to-end platform for developing, training, and deploying AI systems. Through our merger with Voltage Park, we combine developer-first software with cost-efficient, large-scale compute. We serve solo researchers, startups, and large enterprises globally. **The Role** We're seeking an experienced Senior Research Engineer to advance how large language models are trained, fine-tuned, evaluated, and deployed. You'll work across model training, post-training, PyTorch, distributed systems, and AI systems engineering to improve model quality, training efficiency, and developer productivity. **What You'll Do** • Design, build, and optimize training and post-training pipelines for LLMs • Improve model quality through SFT, continued pretraining, preference optimization, RLHF, and evaluation • Build and improve PyTorch-based training infrastructure and developer workflows • Optimize distributed training across multi-GPU environments • Investigate challenging model training issues and performance bottlenecks • Design evaluation methodologies and guide model improvements through experimentation • Collaborate with customers to understand real-world workloads • Partner with research, infrastructure, and platform engineering teams • Contribute to open-source projects **Required Qualifications** • Significant experience training, fine-tuning, and optimizing transformer-based LLMs using PyTorch • Experience with modern LLM techniques (continued pretraining, SFT, RLHF, DPO, PPO, GRPO, reward modeling) • Strong understanding of distributed training and multi-GPU systems • Strong software engineering fundamentals and production-quality Python experience • Experience designing experiments and debugging complex training issues • Excellent communication and collaboration skills • Comfortable in fast-moving, ambiguous environments • Master's degree, PhD, or equivalent industry experience in ML/AI/CS **Ideal Experience** DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, CUDA, Triton, vLLM, GPU optimization, open-source contributions, or startup experience. **Compensation & Benefits** 💰 **Salary:** $165,000 – $310,000 USD ✨ **Benefits Include:** • Comprehensive health coverage (medical, dental, vision) • Meaningful equity (RSUs) • 401(k) matching (U.S.) / pension contributions (U.K.) • Unlimited PTO and company holidays • 2-week company winter break • Paid parental & family leave **Work Arrangement** Hybrid: Minimum 2 in-office days/week in San Francisco, Seattle, NYC, or London. Fully remote considered for candidates outside office hubs.

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord