Sr. Member of Technical Staff
Cerebras · Sunnyvale, CA
About this role
## Sr. Member of Technical Staff **About Cerebras** Cerebras Systems builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (over 10x faster than GPU-based hyperscale cloud inference services). Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy 750MW of scale. **About the Role** Seeking a **Sr. Member of Technical Staff** to design and develop software features that support **system resiliency** and **high availability** across distributed environments. You’ll help build and maintain **scalable AI inference services**, improve reliability through automation, develop cloud deployment workflows, and collaborate across engineering teams. --- ## Responsibilities - Design and develop software features for **system resiliency** and **high availability**, including automated recovery and fault-tolerant architecture. - Build and maintain **cloud-based deployment workflows** for AI inference software on **AWS** to support low-latency, scalable performance. - Develop **Python-based scripts and APIs** to streamline data preprocessing, inference execution, and post-processing for real-time inference. - Use **parallel programming** (multi-threading, async processing) to maximize resource efficiency on AWS compute. - Develop components for **visualization and analysis** of system performance metrics to improve monitoring and usability. - Develop inference software in **Docker** and define **Kubernetes orchestration** strategies for reliability and efficient scaling. - Create automated scripts to detect and mitigate common failure modes. - Debug issues across **model deployment**, **container orchestration**, and **networking configurations**; document steps to reproduce and root-cause defects. - Triage and resolve defects by analyzing **logs, metrics, and distributed traces** (e.g., CloudWatch, Grafana, custom Python tools). - Partner with Product Management and UX to define requirements for inference service interfaces (configuration, monitoring, event logging). - Author detailed technical documentation for infrastructure configurations, inference workflows, and APIs. - Track defects, enhancements, and release notes using **Jira** and **Git**. --- ## Skills & Qualifications **Minimum Requirements** - Master’s degree (or foreign equivalent) in Computer Science or related field. - 18 months of experience in roles such as Information Security Analyst, Software Engineer, Sr. Member of Technical Staff, IT Senior Applications Engineer, or related occupation. **Required Skills** - **Infrastructure-as-Code & automation:** Terraform, AWS CloudFormation, AWS CDK, Ansible - **Containerization & orchestration:** Docker, Kubernetes, AWS EKS, AWS ECS, AWS Fargate, Helm - **Compute & serverless:** AWS EC2, AWS Lambda, Auto Scaling Groups - **Monitoring/logging/tracing:** AWS CloudWatch, AWS X-Ray, ELK, Prometheus, Grafana - **Programming:** Python, Node.js, JavaScript, Flask - **Data storage/caching:** PostgreSQL, Redis, NFS - **CI/CD & version control:** Jenkins, Git --- ## Why Join Cerebras Build a breakthrough AI platform beyond GPU constraints, publish/open-source cutting-edge AI research, work on one of the fastest AI supercomputers, enjoy job stability with startup vitality, and join a simple, non-corporate culture that respects individual beliefs. **Apply:** https://www.cerebras.ai/join-us **Pri
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.