AI Systems Research and Development Engineer – LLM Inference Systems & Optimization
Snowflake · US-WA-Bellevue
About this role
## About the Role At Snowflake, we’re building the era of the **agentic enterprise**—and we’re looking for **AI-native systems thinkers** to help advance the state of the art in **LLM inference systems and optimization**. You’ll work on next-generation, high-performance inference systems—optimizing not only **speed and efficiency**, but also how quickly inference can **adapt to new models, architectures, hardware, and workloads**. ## What You’ll Do - **Design and develop** high-performance LLM inference systems across distributed serving, runtime systems, GPU execution, and performance-critical kernels. - Create **novel techniques** to improve **latency, generation speed, throughput, memory efficiency, scalability, and cost**. - Explore advanced inference methods such as: - **Speculative & parallel decoding** - **Prefill/decode disaggregation** - **Adaptive parallelism** - **Continuous batching & scheduling** - **KV-cache optimization** - **Quantization** - **Communication optimization** - Build **adaptive inference systems** that automatically optimize execution for new model architectures, hardware, workload characteristics, and deployment environments. - Apply **AI-driven / AI-native systems engineering** (automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning). - Identify high-impact performance problems, form hypotheses, prototype solutions, and drive ideas from **research to production**. - Design distributed inference strategies across GPUs and nodes (e.g., **tensor, sequence, pipeline, data, and expert parallelism**). - Develop approaches for **multi-model serving**, dynamic resource management, model loading/swapping, and workload-aware scheduling. - Analyze and optimize **GPU kernels and operators** for attention, MoE, communication, and other performance-critical components. - Explore **model-system co-design** to unlock substantially more efficient inference. - Profile and benchmark end-to-end workloads to find bottlenecks across compute, memory, communication, networking, scheduling, and model execution. - Collaborate with model researchers, infrastructure teams, and product teams to deploy innovations in production. - **Open-source and publish** innovations via technical blogs and top-tier systems/ML conferences. ## Recent Work / Innovations (Examples) - **Arctic Inference** (open-source inference system) - **Shift Parallelism** (dynamically adapts parallelism to workload characteristics) - **SwiftKV** (reduces redundant prefill computation) - **Arctic Speculator** & **SuffixDecoding** (fast speculative decoding) - **Jacobi Forcing** (causal parallel decoding) - **Semi-Persistence** (fast model swapping and dynamic multi-model serving) ## What You Bring (Requirements) - Bachelor’s degree in **Computer Science, Electrical Engineering, or related** (Master’s/PhD preferred). - **5+ years** experience in one or more of: - LLM inference systems - Distributed AI systems - GPU systems - High-performance computing - Strong understanding of modern **LLM inference architectures** and performance tradeoffs. - Hands-on experience with inference/serving frameworks such as **vLLM, SGLang, TensorRT-LLM**, or similar. - Experience designing/extending/optimizing inference runtimes (scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, disaggregated serving).
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.