Principal Engineer, Distributed Systems
CoreWeave · New York, NY / Sunnyvale, CA
About this role
**CoreWeave — Principal Engineer, Distributed Systems** ## About the role We’re looking for a **Principal Engineer** to provide technical leadership across **Security Products**. This is a senior **individual-contributor** role where you’ll define architecture, guide execution across multiple teams, and solve complex **distributed-systems** problems in **security-critical infrastructure**. You’ll help design systems that meet demanding production requirements, including **four nines availability**, **high scalability**, **consistently low latency**, strong **security boundaries**, and **safe behavior under partial failure** in a **multi-region** setup. Your work spans high-scale **authorization**, **anomaly detection**, **bot defense**, **Security Token Service (STS)**, an **API authentication gateway**, and the shared infrastructure needed to operate these capabilities reliably across regions and environments. ## What you will do - Establish architectural approaches for distributed systems (multi-region operation, service placement, failover, replication, traffic management, disaster recovery, and regional independence) - Drive decisions around consistency models, caching, invalidation/propagation, revocation, idempotency, concurrency, and ordering - Design for data residency, tenant isolation, trust boundaries, blast-radius reduction, and controlled handling of sensitive security data - Define fault-tolerance strategies for dependencies, networks, regions, storage systems, and control-plane components (graceful degradation + safe recovery) - Design systems to meet **four-nines availability** while maintaining predictable low latency and high throughput under normal load, spikes, and partial failures - Improve reliability and operability via **SLOs**, error budgets, metrics/logs/traces, audit events, alerting, incident response, and post-incident learning - Guide teams through architecture/design reviews, implementation tradeoffs, capacity planning, load testing, performance analysis, and production readiness - Mentor senior/staff engineers, develop technical talent, and raise engineering quality across the organization - Communicate architecture, tradeoffs, risks, and recommendations clearly to technical and executive audiences - Lead high-scale authorization system design (policy authoring, evaluation, enforcement, auditability, and integration across services/tenants) - Provide technical direction for STS (issuance, validation, lifecycle management, trust relationships, key rotation, revocation, secure service-to-service access) - Guide design/implementation of an API Authentication Gateway for consistent, secure, observable authentication and authorization - Set technical standards for API design, service contracts, threat modeling, cryptographic key management, secrets handling, observability, and operational readiness - Partner with service teams to make capabilities easy to adopt (interfaces, SDKs, reference implementations, documentation, integration patterns) ## What you bring - **12+ years** designing/building production software, including substantial distributed-systems and platform infrastructure experience - Track record as a principal/staff-plus/distinguished (or equivalent) technical leader across teams or a broad domain - Deep expertise in distributed-systems design (scalability, availability, latency, consistency, concurrency, partition tolerance, caching, replication, failure recovery) - Experience designing/ope
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.