ML Systems Integration Engineer
Cerebras · Sunnyvale, CA
About this role
## ML Systems Integration Engineer **About Cerebras** Cerebras Systems builds large-scale AI hardware designed to deliver industry-leading training and inference speeds—enabling real-time iteration and more agentic computation. Cerebras works with leading model labs, global enterprises, and AI-native startups. (See: https://openai.com/index/cerebras-partnership/) --- ## Responsibilities - Participate in bring-up of next-generation AI hardware systems and supporting software infrastructure - Debug complex system-level issues spanning hardware and software interactions - Investigate system bring-up failures and identify root causes using logs, telemetry, and diagnostic tools - Build automation frameworks and internal tooling to improve system validation and debugging workflows - Develop software to test, validate, and stress distributed hardware systems during development and production - Collaborate closely with hardware engineers to isolate and resolve system integration issues - Improve system observability by building tools that surface failures quickly and accelerate debugging - Reproduce, triage, and diagnose difficult issues during early hardware deployment - Support validation and qualification of new hardware generations toward production readiness - Continuously improve internal engineering workflows for debugging, testing, and automation --- ## Skills & Qualifications - BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related technical field - Strong programming skills in **Python and/or C++** - Excellent debugging and problem-solving skills; methodical investigation of complex issues - Solid understanding of operating systems fundamentals (processes, threads, memory management, concurrency, IPC) - Experience working in **Linux** development environments - Understanding of computer architecture and hardware/software interactions - Strong analytical thinking to break down complex failures into actionable root causes - Ability to collaborate effectively across multiple engineering teams - Strong communication skills and comfort working on ambiguous technical problems --- ## Preferred Qualifications - Experience building automation frameworks, internal tooling, or test infrastructure - Familiarity with distributed systems concepts - Experience debugging large-scale systems or complex infrastructure environments - Understanding of networking fundamentals and communication between distributed systems - Experience with hardware-adjacent software / system integration environments - Familiarity with performance analysis, system telemetry, and log analysis - Exposure to production systems validation or infrastructure reliability engineering --- ## Why Join Cerebras - Build a breakthrough AI platform beyond GPU constraints - Publish and open source cutting-edge AI research - Work on one of the fastest AI supercomputers in the world - Job stability with startup vitality - A simple, non-corporate work culture Learn more: https://www.cerebras.ai/join-us --- *Equal Opportunity Employer.* Privacy/CCPA: https://www.cerebras.net/privacy/
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.