CronJobs

backend jobs

ML Runtime and Kernel Engineer - Core ML

Cerebras · Sunnyvale, CA

hybridunknownPosted Sep 28, 2026C++PythonPyTorchJAXCUDATriton

Apply on the employer site

About this role

## ML Runtime and Kernel Engineer - Core ML ### About the Role Cerebras Systems builds the world’s largest AI chip (Wafer-Scale Engine), enabling industry-leading training and inference speeds—over 10× faster than GPU-based hyperscale cloud inference services. The Core ML team develops novel machine learning algorithms that leverage the unique capabilities of the Cerebras platform. You’ll bridge the gap between promising research ideas and efficient execution on Cerebras systems. Work across ML frameworks, compilers, runtimes, and low-level kernels to implement new algorithmic capabilities, diagnose performance bottlenecks, and turn research prototypes into robust, high-performance demonstrations. Depending on your background, your work may focus on: - **Runtime capabilities** (token orchestration, scheduling, communication, distributed execution) - **Low-level kernel development** for novel ML operations - **A combination of both** ### Responsibilities - Design and implement **runtime components** and **high-performance kernels** required by novel Core ML algorithms. - Translate **research prototypes** into efficient Cerebras implementations, including reference implementations and GPU comparisons where useful. - Profile and debug performance across the **ML framework, compiler, runtime, communication, and kernel** layers. - Optimize **computation, memory movement, communication, and concurrency** for large-scale training and low-latency inference. - Develop **benchmarks, instrumentation, and automated tests** to validate functionality, performance, and numerical correctness. - Collaborate with Core ML researchers and engineering teams (compiler, runtime, kernels, inference) to deliver end-to-end capabilities. - Contribute to **software architecture and roadmap** decisions by identifying recurring limitations and high-leverage platform improvements. ### Skills & Qualifications - Bachelor’s, Master’s, PhD, or equivalent practical experience in **Computer Science, Computer Engineering, Electrical Engineering,** or related fields. - Experience developing **high-performance systems software**, **ML systems**, **runtimes**, **compilers**, or **computational kernels**. - Strong programming skills in **C++** and **Python**. - Solid understanding of **parallel programming**, **memory management**, **concurrency**, **data structures**, and **performance optimization**. - Proven ability to debug and profile complex software across multiple system layers. - Familiarity with modern ML architectures/frameworks such as **PyTorch** or **JAX**. - Ability to work effectively with researchers and translate evolving algorithmic requirements into reliable software. ### Preferred Skills & Qualifications - Experience with **CUDA**, **Triton**, **low-level assembly**, accelerator programming, or a C-like domain-specific language. - Experience with **compiler internals**, **distributed runtimes**, custom hardware interfaces, or **HPC systems**. - Understanding of ML fundamentals and ML systems—reasoning about how algorithmic choices affect accuracy, implementation, and performance. - Familiarity with **LLM training/inference** (attention, KV-cache management, parallel generation, distributed execution). - Experience developing software in industrial or academic research environments where requirements evolve through experimentation. - Contributions to significant **open-source systems**, ML frameworks, compilers, or kernel libraries. ### Why Join Cereb

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord