Principal Machine Learning Engineer
Bjakcareer · United States
About this role
**About the Role** Build proactive, reliable AI for everyday users who aren’t “AI-native.” You’ll own the execution layer that turns research and model capabilities into production-grade, scalable ML systems—covering the full model lifecycle: data, training, evaluation, inference, and deployment. **What You’ll Own** - End-to-end ML systems powering the company (data → training → evaluation → inference → deployment) - Training and fine-tuning pipelines for large models - Evaluation systems to measure capability, robustness, safety, and real-world product performance - High-performance inference systems (latency, GPU utilization, memory, cost, reliability) - Data pipelines for high-quality real-world and synthetic training data - Production infrastructure for deploying, monitoring, and continuously improving models - Close partnership with research and application engineering to drive product improvements - Pragmatic technical trade-offs and rapid iteration based on real-world performance **What We’re Looking For** - Experience building and shipping production ML systems (not just research prototypes) - Strong understanding of large-model training, fine-tuning, evaluation, and inference - Strong software engineering and systems fundamentals - Experience operating ML workloads at meaningful scale (especially GPU-based) - Ability to navigate ambiguous problems independently with strong technical judgment - Bias toward experimentation, measurement, and shipping - High standards for correctness, reliability, and production quality **Outcomes** - Research and models reliably translate into production-ready solutions with clear performance/quality targets - ML pipelines, training loops, and inference systems are stable, efficient, and maintainable - Production issues are detected, debugged, and resolved quickly - Team members are supported and able to deliver high-impact ML work with minimal friction - Model/system iterations are measurable, safe, and improve user experience over time **Tech Stack** - Python - PyTorch / JAX - GPU-based training and inference systems **Ideal Experience** - Built or shipped real ML systems used by people (not just demos) - Comfortable working with large models and understanding their failure modes - Writes strong, production-grade code with a focus on system correctness **How We Work** Small, high-talent-density, hands-on team with broad ownership. Decisions are made quickly, with less process and more building something exceptional. **Interview Process** If there’s a fit, you’ll be scheduled for **3 interviews** (no more than **4**). Interviews are virtual and/or onsite. Expect a prompt decision. *This is an invitation to help bring AI with practical benefits to billions globally.*
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.