CronJobs

backend jobs

Engineering Manager, Cloud Monitoring Services Platform

Crusoe · San Francisco, CA - US

onsiteunknown$215,000–$260,000Posted Sep 25, 2026GoRustJavaNode.jsKubernetesPrometheusClickHouseOpenTelemetry

Apply on the employer site

About this role

## Engineering Manager, Cloud Monitoring Services Platform Crusoe is on a mission to accelerate the abundance of energy and intelligence. As a vertically integrated AI infrastructure company, we own and operate each layer of the stack—from electrons to tokens—to power the world’s most ambitious AI workloads. ### About the Role Crusoe builds cloud infrastructure for AI workloads. **Cloud Monitoring Services** owns observability across Crusoe Cloud: **metrics, logs, alerting, and the telemetry agent** that runs on every node in the fleet. We’re hiring an **Engineering Manager** to lead the **Platform team**—focused on the **time series and log storage systems** and the **query layer** that serves every dashboard, API call, and investigation. This is a **first-line management** role reporting to the Engineering Manager for Cloud Monitoring Services. You’ll manage **4–6 engineers**, growing the team over time, and own how telemetry is **stored, retained, and read back at fleet scale**. This is a **people-first leadership** role with real delivery stakes. ### What You’ll Be Working On - **Grow and develop your team** (4–6 engineers): 1:1s, career growth, performance, and team health. - **Own storage and query**: time series + log storage systems and the query layer on top, including behavior under load and cost. - **Own delivery**: plan and sequence work across a roadmap mixing customer-facing features with storage/query infrastructure; escalate early when plans are at risk. - **Set technical direction**: partner with Staff engineers to ask hard questions and make sound tradeoff calls. - **Keep the operational bar high**: query is on the critical path during fleet issues—availability, performance, and correctness are core. - **Own the oncall rotation** and operational health of the storage/query stack. - **Manage cost and scale together**: retention policy, downsampling, cardinality, and storage tiering. - **Coordinate across team boundaries**: consume outputs from collection/ingestion and serve internal and external customers. - **Shape the team over time**: recruiting support, interview loop, and onboarding. - **Collaborate across functions**: align priorities with product, neighboring infrastructure teams, peer managers, and leadership. ### What You’ll Bring to the Team - **People-first mindset**: you want your engineers to grow; you measure success through your team. - **Technical depth**: 5+ years hands-on backend/distributed systems experience (e.g., databases, storage engines, query engines, streaming pipelines) and familiarity with **Kubernetes**. - **Engineering management experience**: 2+ years managing software engineers directly, including performance cycles and career conversations. - **Coaching experience**: comfort partnering with strong technical leads (senior/staff). - **Track record of shipping**: delivered multi-phase projects against fixed deadlines. - **Operational judgment**: ran teams owning systems customers depend on during incidents. - **Strong communication and judgment**: align technical and non-technical partners; provide clarity on what matters and why. ### Bonus Points - Time series databases / columnar storage at scale (e.g., Prometheus, Thanos, Mimir, VictoriaMetrics, InfluxDB, ClickHouse) - Query engines and query cost control (planning, pushdown, caching, concurrency limits) - Retention, compaction, downsampling, storage tiering for large telemetry/event datasets - Log storage/search systems (e.g., L

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord