Senior Staff Cloud Backend Engineer
Coupang · Seattle, USA
About this role
**Senior Staff Cloud Backend Engineer** **Company Introduction** We exist to wow our customers. Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. Coupang is one of the fastest-growing e-commerce companies with a reputation for being a dominant and reliable force in South Korean commerce. We’re proud to have the best of both worlds—startup culture with the resources of a large global public company. At Coupang, you’ll see yourself, your colleagues, your team, and the company grow every day. --- **Role Overview** - Implement SRE best practices to improve reliability, scalability, and performance of datacenter services. - Develop and maintain automation scripts for infrastructure provisioning, monitoring, and management. - Conduct root cause analysis and post-mortem reviews to prevent recurrence of incidents. --- **Roles and Responsibilities** **Observability and Monitoring** - Design, implement, and maintain observability solutions for datacenter infrastructure. - Develop, deploy, and maintain operational and reliability components of a large-scale Observability and Telemetry collection platform (performance at scale, real-time monitoring, logging, and alerting). - Participate in the full lifecycle of services—from inception and design to deployment, operation, and refinement. - Develop and optimize monitoring systems to ensure high availability and performance. - Create and manage dashboards, alerts, and reports for system health and performance visibility. **Performance Optimization** - Analyze and optimize performance of datacenter systems and applications. - Implement best practices for resource utilization and efficiency. **Collaboration** - Work closely with other engineering teams to meet observability and reliability requirements. - Collaborate with hardware and software vendors to evaluate and integrate new technologies. **Security and Compliance** - Ensure observability and reliability solutions comply with security policies and industry standards. - Implement and maintain security measures to protect data and infrastructure. **Troubleshooting and Support** - Support observability and reliability-related issues, including debugging and resolving hardware and software problems. - Develop and maintain documentation for troubleshooting procedures and best practices. **Continuous Improvement** - Stay updated on the latest advancements in observability and SRE technologies and integrate them into the infrastructure. - Continuously improve reliability, scalability, and performance of datacenter services. --- **Qualifications** - Bachelor’s degree in Computer Science, Electrical Engineering, Math, or a closely related field - 8+ years of experience in backend software development - Experience working in cloud environments, particularly AWS - Demonstrated experience building and maintaining highly available, distributed systems --- **Preferred Qualifications** - Proficiency in observability tools/technologies (e.g., Prometheus, Grafana, ELK Stack) - Experience with SRE practices and tools (e.g., Kubernetes, Docker, Terraform) - Strong programming and scripting skills (e.g., Python, Go, Bash) - Familiarity with cloud platforms (AWS, Azure, GCP) and their observability/reliability services - Strong problem-solving skills and attention to detail - Excellent communication and collaboration skills - Abi
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.