Senior Software Engineer, AI/ML Platform
Agility Robotics · Remote
About this role
**About the Role** Join the team building the machine learning platform to power fleet-scale humanoid robotics. As a Senior Engineer on the ML Infrastructure and Platform group, you will help **architect** and **build** foundational infrastructure for AI/ML operations at Agility—covering the platform layer for data collection & processing, training, sim/real evaluation, and model management & observability. Your work will empower AI teams across perception, controls, skills, and innovation to build and deploy next-generation robot foundation models and end-to-end policies. **Key Responsibilities** - **Execution and Technical Ownership** - Contribute to the design and implementation of the ML platform orchestrating the end-to-end AI flywheel: data processing, training, evaluation, and deployment - Develop reliable workflows across cloud compute, Kubernetes, and continuous automation - Build core infrastructure components such as the model registry, feature store, and experiment tracking tooling - Own developer-facing APIs and CLI tools to make ML workflows simple and reproducible - Implement CI/CD for ML to enable continuous retraining, automated testing, and seamless model delivery to production - **Collaboration** - Work closely with the Staff ML Infra Engineer and cross-functional stakeholders (AI researchers and robotics engineers) to translate requirements into scalable systems - Partner with data platform engineers to integrate ML orchestration and metadata tracking with existing data lake and pipelines - **Engineering Excellence, Growth and Impact** - Apply MLOps best practices: reproducibility, lineage, rollback, monitoring, and governance - Mentor junior engineers and influence the broader cloud platform roadmap - Contribute to internal discussions on platform architecture, reliability, and scalability **What We’re Aiming For (MLOps Level 2)** - Version-controlled ML pipelines (data, code, and config) - Automated and reproducible model training and evaluation - Continuous integration and delivery for ML workflows - Centralized experiment tracking and performance visualization - Standardized model packaging and deployment to production - Monitoring of models post-deployment **Required Qualifications** - 5+ years of software engineering experience, with at least 2+ years building ML infrastructure, data platforms, or MLOps systems in production - Experience building/maintaining components of modern ML platforms (e.g., experiment tracking, model registries, training pipelines, deployment systems) - Familiarity with orchestration and tracking tools (MLflow, Weights & Biases, Airflow, Kubeflow, etc.) - Proficiency with cloud-native platforms (AWS, GCP, or Azure), containers, and IaC (e.g., CDK, Terraform) - Hands-on experience with processing or modeling multimodal data (sensor logs, camera streams, behavior traces, etc.) - Comfortable collaborating cross-functionally to ship infrastructure used by others **Bonus Qualifications** - Experience with robotics, autonomous vehicles, drones, or embedded ML - Contributions to open-source ML infrastructure or MLOps tooling (a plus) **Why This Role?** - **Build from the start:** Help shape the ML platform layer as it’s being defined - **High impact:** Enable faster, safer, more intelligent robotic behaviors at scale - **Technical frontier:** Work on enabling the next frontier of AI in real production settings - **Remote-friendly:** Strong engineering culture with a f
Listing freshness
CronJobs last confirmed this listing 9h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.