Staff Backline Engineer – ML/AI
Databricks · Bellevue, Washington; San Francisco, California
About this role
**Staff Backline Engineer – ML/AI** **About the Team** The Backline Engineering Team serves as the critical bridge between Frontline Support and Engineering. You’ll handle complex technical issues and escalations across the Data and AI ecosystem, with a strong focus on customer success through deep technical expertise, proactive issue resolution, and continuous platform improvements. **What You’ll Do** - Serve as a senior escalation point for complex ML/AI issues across model training, inference, Model Serving, MLflow, and Feature Engineering (with knowledge of Spark, Delta Lake, and distributed workloads). - Perform deep technical investigations using logs, traces, metrics, profiling, configuration, source code, and customer workloads to identify root cause. - Reproduce customer issues via hands-on experimentation, Python/Spark development, workload construction, configuration changes, and performance analysis. - Troubleshoot model training/inference failures, performance degradation, resource utilization, memory/CPU/GPU issues, distributed execution problems, and deployment/runtime failures. - Partner with Engineering and Product to drive difficult issues to resolution and influence product improvements. - Identify recurring failure patterns and turn them into better diagnostics, documentation, tooling, automation, and Claude skill capabilities. - Mentor engineers and raise the technical troubleshooting capabilities of the broader Support organization. - Act as a technical SME for ML/AI platform areas and contribute to cross-functional initiatives with global impact. **What You’ll Bring** - Deep troubleshooting experience with distributed ML/AI systems; ability to debug across application code, frameworks, infrastructure, and the Databricks platform. - Strong hands-on Python experience; ability to build, modify, and debug ML workloads using PyTorch, TensorFlow, or Scikit-Learn. - Strong understanding of Databricks ML/AI technologies (MLflow, Model Serving, Feature Engineering, Spark MLlib, and model lifecycle management). - Strong Apache Spark knowledge (DataFrames, query execution, distributed computing, memory management, shuffles, and performance optimization). - Experience troubleshooting training/inference performance (CPU/GPU utilization, memory issues, data bottlenecks, concurrency, latency, distributed execution). - Experience with ML deployment and infrastructure such as Kubernetes, cloud ML platforms, CI/CD, model monitoring, and production ML systems. - Ability to read and reason about code, logs, stack traces, metrics, traces, execution plans, and system behavior. - Demonstrated ability to independently own ambiguous, high-impact technical problems and drive them from **symptom → investigation → root cause → resolution**. - Strong technical communication skills and ability to influence Engineering, Product, Support, and customers. **What We Look For** - Customer-obsessed candidates with **10+ years of relevant experience**. - Deep expertise in one of the following tracks (excellence in one area rather than proficiency in all): - **Data Engineering Track:** Large-scale big data solutions and ETL pipelines using Spark, Delta Lake, or Hive; troubleshooting failures and performance; strong Python/SQL/Scala. - **Product Supportability Track:** Distributed system internals; code-level root-cause analysis and profiling using metrics and heap/thread dumps (Java/Scala/Python); bug fixes and mentoring. - **AI Track:** Larg
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.