Staff Software Engineer, Replication Foundations
Temporal · United States (Remote)
About this role
**Staff Software Engineer — Replication Foundations** **Role Summary** We’re hiring a **Staff Software Engineer** to join the **Replication Foundations** team within **Temporal’s Cloud Global Services (CGS)** organization. The team owns and evolves Temporal’s core **replication stack**—the distributed systems backbone behind key Temporal Cloud capabilities such as **High Availability namespaces**, **cross-cluster/cross-region failover**, and **migration products** that help customers move workloads between **self-hosted Temporal** and **Temporal Cloud**. You’ll help set the technical direction for Temporal’s distributed replication systems, leading correctness-critical initiatives across architecture, design, implementation, rollout, and operations. **What You’ll Do** - Set the technical direction and evolve the architecture of Temporal’s **OSS replication stack**, from problem definition through rollout and operational support. - Lead the design and implementation of replication protocols that power: - **High Availability namespaces** - **Cross-cluster and cross-region replication** - **Migration between Temporal clusters**, including **cloud-to-self-hosted** and **cloud-to-cloud** scenarios - Drive scalability and reliability initiatives such as: - **Multi-cell namespaces** - Enabling a namespace to span **multiple clusters** - Improving **load distribution** and handling **hot spots** - Define and communicate system-level guarantees, including: - **Consistency models**, **ordering**, **idempotency** - **Failure recovery** and **performance** - **Operational behavior** - Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities supporting current and future cloud products. - Partner with **Cloud Enablement**, **CGS**, **Product**, and other engineering teams to align OSS replication foundations with customer and product needs. - Lead design reviews, raise the quality of implementation and testing practices, mentor engineers, and provide technical guidance across the organization. - Lead or contribute to debugging complex production issues, including incident response and follow-up improvements related to replication and core system behavior. **What You’ll Bring** - A track record of designing and delivering complex production **distributed systems** where **correctness, availability, and performance** are critical. - Deep understanding of distributed systems fundamentals: **replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery**. - Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces. - Experience debugging complex production issues (e.g., concurrency bugs, data inconsistencies, partial failures, performance bottlenecks). - Proficiency writing production-quality concurrent code in **Go** (Java, C++, or similar systems languages are a plus). - Strong written and verbal communication skills—able to explain complex designs and trade-offs to technical and cross-functional audiences. - Demonstrated ability to influence technical direction across teams, build alignment without direct authority, and mentor engineers. - A thoughtful, curious approach to understanding system behavior under **load, failure, and changing workloads**. **Nice to Have** - Experience designing or maintaining replication protocols or **data-plane infrastructure**. - Experien
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.