CronJobs

devops-sre jobs

Sr. Member of Technical Staff

Cerebras · Sunnyvale, CA

onsiteseniorPosted May 8, 2026PythonNode.jsJavaScriptFlaskAWSKubernetesDockerTerraform

Apply on the employer site

About this role

## Sr. Member of Technical Staff **About Cerebras** Cerebras Systems builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (over 10x faster than GPU-based hyperscale cloud inference services). Cerebras works with leading model labs, global enterprises, and AI-native startups, including a recent multi-year partnership with OpenAI to deploy 750MW of scale. **About the Role** Seeking a **Sr. Member of Technical Staff** to design and develop software features that support **system resiliency** and **high availability** across distributed environments. You’ll help build and maintain **scalable AI inference services**, improve reliability through automation, develop cloud deployment workflows, and collaborate across engineering teams. --- ## Responsibilities - Design and develop software features for **system resiliency** and **high availability**, including automated recovery and fault-tolerant architecture. - Build and maintain **cloud-based deployment workflows** for AI inference software on **AWS** to support low-latency, scalable performance. - Develop **Python-based scripts and APIs** to streamline data preprocessing, inference execution, and post-processing for real-time inference. - Use **parallel programming** (multi-threading, async processing) to maximize resource efficiency on AWS compute. - Develop components for **visualization and analysis** of system performance metrics to improve monitoring and usability. - Develop inference software in **Docker** and define **Kubernetes orchestration** strategies for reliability and efficient scaling. - Create automated scripts to detect and mitigate common failure modes. - Debug issues across **model deployment**, **container orchestration**, and **networking configurations**; document steps to reproduce and root-cause defects. - Triage and resolve defects by analyzing **logs, metrics, and distributed traces** (e.g., CloudWatch, Grafana, custom Python tools). - Partner with Product Management and UX to define requirements for inference service interfaces (configuration, monitoring, event logging). - Author detailed technical documentation for infrastructure configurations, inference workflows, and APIs. - Track defects, enhancements, and release notes using **Jira** and **Git**. --- ## Skills & Qualifications **Minimum Requirements** - Master’s degree (or foreign equivalent) in Computer Science or related field. - 18 months of experience in roles such as Information Security Analyst, Software Engineer, Sr. Member of Technical Staff, IT Senior Applications Engineer, or related occupation. **Required Skills** - **Infrastructure-as-Code & automation:** Terraform, AWS CloudFormation, AWS CDK, Ansible - **Containerization & orchestration:** Docker, Kubernetes, AWS EKS, AWS ECS, AWS Fargate, Helm - **Compute & serverless:** AWS EC2, AWS Lambda, Auto Scaling Groups - **Monitoring/logging/tracing:** AWS CloudWatch, AWS X-Ray, ELK, Prometheus, Grafana - **Programming:** Python, Node.js, JavaScript, Flask - **Data storage/caching:** PostgreSQL, Redis, NFS - **CI/CD & version control:** Jenkins, Git --- ## Why Join Cerebras Build a breakthrough AI platform beyond GPU constraints, publish/open-source cutting-edge AI research, work on one of the fastest AI supercomputers, enjoy job stability with startup vitality, and join a simple, non-corporate culture that respects individual beliefs. **Apply:** https://www.cerebras.ai/join-us **Pri

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord