AI Inference Core - Infrastructure SW Engineer
Cerebras · Sunnyvale, CA
About this role
## AI Inference Core - Infrastructure SW Engineer ### About Cerebras Cerebras Systems builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (over 10x faster than GPU-based hyperscale cloud inference). This performance increase is transforming AI application experiences by enabling real-time iteration and additional agentic computation. ### About the Team The Core Infrastructure team builds the software systems that power engineering workflows across Cerebras. You’ll work on infrastructure that coordinates complex work across machines, clusters, development environments, and hardware systems—covering orchestration frameworks, execution engines, scheduling, test infrastructure, developer tools, and reusable software platforms. ### About the Role We’re hiring a Software Engineer to design and build the core software behind Cerebras engineering infrastructure. You’ll work on Python frameworks, orchestration systems, distributed execution, scheduling, test infrastructure, and developer tooling. This role is a great fit if you enjoy: - Reading unfamiliar code and understanding how systems fit together - Debugging difficult problems - Improving underlying design (not just applying one-off fixes) ### Responsibilities - Design, develop, test, and maintain Python frameworks and services for orchestrating engineering workflows across machines and clusters - Build reusable abstractions for scheduling, distributed execution, resource management, test execution, workflow planning, and failure recovery - Define clear APIs, module boundaries, extension points, and data models for long-term maintainability - Reason about concurrency, async execution, multiprocessing, state management, retries, idempotency, cancellation, and partial failures - Debug complex issues across Python, OS/processes, filesystems, networking, remote machines, and distributed services - Write high-quality automated tests and documentation for widely reused infrastructure - Partner with platform, CI, release, quality, ML systems, and product engineering teams to translate requirements into scalable designs ### Skills & Qualifications - 3+ years of professional software engineering experience - Strong Python proficiency and understanding of runtime behavior - Experience designing maintainable systems, libraries/frameworks, backend services, or developer-facing APIs - Strong architecture judgment (abstraction boundaries, design patterns, extensibility) - Concurrency fundamentals (processes, threads, async execution, synchronization, shared state) - OS fundamentals (processes, signals, filesystems, resource management) - Distributed systems fundamentals (retries, timeouts, idempotency, partial failure, coordination, eventual consistency) - Strong debugging and independent problem-solving skills ### Preferred Qualifications - Python concurrency technologies (asyncio, multiprocessing, concurrent futures, event-driven systems) - Experience building orchestration engines/workflow systems/schedulers/distributed job runners/control-plane software - Test infrastructure experience (e.g., pytest extensions) - Familiarity with CI/build/release/developer productivity tooling - Kubernetes, containers, cluster schedulers, or remote execution systems - BS/MS in CS or related field (or equivalent practical experience) ### Why Join Cerebras People at Cerebras say they joined for reasons including: 1. Building a breakthrough AI platform
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.