CronJobs

ai-ml jobs

Staff Backline Engineer – ML/AI

Databricks · Bellevue, Washington; San Francisco, California

remotesenior$170,400–$255,600Posted Sep 14, 2026PythonApache SparkDelta LakeMLflowPyTorchTensorFlowKubernetesSQL

Apply on the employer site

About this role

**Staff Backline Engineer – ML/AI** **About the Team** The Backline Engineering Team serves as the critical bridge between Frontline Support and Engineering. You’ll handle complex technical issues and escalations across the Data and AI ecosystem, with a strong focus on customer success through deep technical expertise, proactive issue resolution, and continuous platform improvements. **What You’ll Do** - Serve as a senior escalation point for complex ML/AI issues across model training, inference, Model Serving, MLflow, and Feature Engineering (with knowledge of Spark, Delta Lake, and distributed workloads). - Perform deep technical investigations using logs, traces, metrics, profiling, configuration, source code, and customer workloads to identify root cause. - Reproduce customer issues via hands-on experimentation, Python/Spark development, workload construction, configuration changes, and performance analysis. - Troubleshoot model training/inference failures, performance degradation, resource utilization, memory/CPU/GPU issues, distributed execution problems, and deployment/runtime failures. - Partner with Engineering and Product to drive difficult issues to resolution and influence product improvements. - Identify recurring failure patterns and turn them into better diagnostics, documentation, tooling, automation, and Claude skill capabilities. - Mentor engineers and raise the technical troubleshooting capabilities of the broader Support organization. - Act as a technical SME for ML/AI platform areas and contribute to cross-functional initiatives with global impact. **What You’ll Bring** - Deep troubleshooting experience with distributed ML/AI systems; ability to debug across application code, frameworks, infrastructure, and the Databricks platform. - Strong hands-on Python experience; ability to build, modify, and debug ML workloads using PyTorch, TensorFlow, or Scikit-Learn. - Strong understanding of Databricks ML/AI technologies (MLflow, Model Serving, Feature Engineering, Spark MLlib, and model lifecycle management). - Strong Apache Spark knowledge (DataFrames, query execution, distributed computing, memory management, shuffles, and performance optimization). - Experience troubleshooting training/inference performance (CPU/GPU utilization, memory issues, data bottlenecks, concurrency, latency, distributed execution). - Experience with ML deployment and infrastructure such as Kubernetes, cloud ML platforms, CI/CD, model monitoring, and production ML systems. - Ability to read and reason about code, logs, stack traces, metrics, traces, execution plans, and system behavior. - Demonstrated ability to independently own ambiguous, high-impact technical problems and drive them from **symptom → investigation → root cause → resolution**. - Strong technical communication skills and ability to influence Engineering, Product, Support, and customers. **What We Look For** - Customer-obsessed candidates with **10+ years of relevant experience**. - Deep expertise in one of the following tracks (excellence in one area rather than proficiency in all): - **Data Engineering Track:** Large-scale big data solutions and ETL pipelines using Spark, Delta Lake, or Hive; troubleshooting failures and performance; strong Python/SQL/Scala. - **Product Supportability Track:** Distributed system internals; code-level root-cause analysis and profiling using metrics and heap/thread dumps (Java/Scala/Python); bug fixes and mentoring. - **AI Track:** Larg

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord