CronJobs

devops-sre jobs

Software Engineer II Platform Data Reliability

Sony Interactive Entertainment Global · United States, San Mateo, CA

hybridmid$150,100–$150,100Posted Sep 12, 2026GoTerraformAnsibleKubernetesAWSGCPKafkaRedis

Apply on the employer site

About this role

**Why Sony Interactive Entertainment?** Sony Interactive Entertainment (SIE) isn’t just the Best Place to Play—it’s also the Best Place to Work. As the company behind the PlayStation brand, SIE delivers cutting-edge hardware and network services to more than 100 million people, and creates experiences for beloved PlayStation IP. --- ## **Software Engineer II — Platform Data Reliability & Automation** Ready to level up your career? Join PlayStation as a **Software Engineer II** focused on **Platform Data Reliability & Automation** and help build reliable, scalable data platforms for millions of players worldwide. You’ll improve reliability and automation across **NoSQL, streaming, and caching** services in **AWS and GCP**, and build operational tooling and observability for technologies such as **Cassandra, Aerospike, Kafka, and Redis**. --- ## **Role Overview** You’ll build, automate, and operate scalable data platforms using **Infrastructure as Code (IaC)** and cloud technologies. Working with senior engineers, platform teams, and product teams, you’ll help reduce manual work, improve uptime, and make data services safer and easier for engineering teams to use. --- ## **Responsibilities** - Develop, maintain, and improve **Infrastructure as Code** and configuration-management automation (e.g., **Terraform**, **Ansible**) to provision, configure, monitor, scale, and manage NoSQL/streaming/caching platforms. - Build automation for **repeatable and reliable deployments** across cloud and hybrid environments. - Improve **reliability, availability, scalability, performance, and resiliency** of platform data services. - Define, measure, and improve **service-level indicators (SLIs)**, **service-level objectives (SLOs)**, and **error budgets**. - Automate operational activities such as **scaling, failover, backup, recovery, upgrades, and routine maintenance**. - Build and enhance **observability** (metrics, logging, tracing, dashboards, alerts). - Troubleshoot issues affecting **Cassandra, Aerospike, Kafka/MSK, Redis**, and related services. - Participate in **on-call** and incident response, including root-cause analysis and permanent fixes. - Write reliable, maintainable, well-tested **Go** code for automation and operational tooling. - Collaborate across engineering, platform, security, and operations teams to deliver reliable data services. - Create and maintain operational documentation, **runbooks**, and automation playbooks. - Participate in code reviews, technical design discussions, and continuous improvement. - Explore practical uses of **AI-assisted automation**, anomaly detection, automated remediation, and developer productivity tooling where appropriate. --- ## **Skills and Qualifications** - Bachelor’s or Master’s degree in Computer Science (or related field), or equivalent practical experience. - **3+ years** in software engineering, database reliability engineering, site reliability engineering, platform engineering, or related work. - Production experience developing in **Go** (idiomatic code, testing, concurrency, error handling, maintainability). - Hands-on experience with **IaC/configuration management** (e.g., **Terraform**, **Ansible**). - Experience deploying/operating workloads on **Kubernetes**. - Experience with **AWS or GCP**, including familiarity with managed services such as **MSK, DynamoDB, ElastiCache, Memorystore**, or equivalents. - Working knowledge of NoSQL/caching/streaming technologies (e.g., **

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord