Software Engineer, Inference (AI Data Engineering)
SpaceX · Palo Alto, CA
About this role
**Software Engineer, Inference (AI Data Engineering)** **About the Role** Join SpaceX's application software team as a Software Engineer focused on AI inference and data engineering. You'll design and optimize large-scale model serving systems end-to-end, from distributed infrastructure to deep low-level optimizations. Work on systems delivering reliable, high-throughput inference powering SpaceX's mission-critical applications. **Key Responsibilities** • Develop highly reliable, high-throughput inference systems serving AI models across SpaceX • Architect scalable distributed infrastructure including load balancing, auto-scaling, batch scheduling, and continuous batching • Optimize latency and throughput with GPU kernel work, quantization, speculative decoding, and acceleration techniques • Build high-concurrency serving systems with 100% uptime and low tail latency • Own end-to-end components: request routing, SDK development, rate limiting, and scaling • Benchmark and accelerate inference engines (SGLang, vLLM, TensorRT-LLM) • Develop tracing and debugging tools across the full stack • Create robust CI/CD infrastructure for seamless deployment • Collaborate across SpaceX AI teams **Basic Qualifications** • Bachelor's in CS, engineering, math, or scientific discipline (or 2+ years professional software experience) • Experience designing and maintaining reliable, horizontally scalable distributed systems • 1+ years full stack or backend development with production systems • 1+ years experience with Rust or C++ **Preferred Skills** • LLM inference engines and serving frameworks (SGLang, vLLM, Triton, TensorRT-LLM) • Deep systems programming: GPU kernels, code generation, batching, caching, quantization • Large-scale, high-concurrency production serving systems • Service observability and reliability best practices • Docker, Kubernetes, containerized applications • gRPC expertise • Python, Go, or similar languages • CI/CD, build systems, and monitoring • Performance profiling and optimization **Additional Requirements** • Onsite in Palo Alto (no remote/hybrid) • May work extended hours/weekends based on launch cadence • ITAR eligible: U.S. citizen, national, green card holder, refugee, or asylee status required **Compensation** Level 1: $135,000–$175,000 | Level 2: $155,000–$210,000 Plus: stock options, bonuses, comprehensive benefits, 401(k), 3 weeks PTO, 10+ paid holidays
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.