Senior Data Infrastructure Engineer
Aircallioinc · San Francisco Office
About this role
**Aircall — Senior Data Infrastructure Engineer** Aircall is an AI-powered customer communications platform used by 22,000+ companies worldwide to drive revenue, resolve issues faster, and scale customer-facing teams. At Aircall, you’ll join a company in motion—ambitious, product-driven, and execution-focused, with visible impact and fast decisions. --- ## About the role Aircall’s Data team is mid-migration: moving off a single Redshift cluster onto an Apache Iceberg lakehouse on S3, with Flink CDC into Kafka for ingestion and dbt-on-Spark via Apache Kyuubi on EKS for transformation. This is a dedicated platform seat. You’ll build high-leverage foundations (data quality automation, schema registry, staging + gated promotion) so analytics engineers, data scientists, and AI agents can move faster. --- ## What you will do - **Build and operate the lakehouse**: Apache Iceberg on S3, table design/maintenance, partitioning + compaction, and migration of remaining Redshift workloads - **Own ingestion end to end**: Flink CDC → Kafka (MSK) → Iceberg, plus Rudderstack, Fivetran, and DMS sources; maintain freshness + reliability SLAs - **Run and evolve compute + orchestration**: Apache Kyuubi on EKS for dbt-spark, Airflow (completing ECS → EKS migration), autoscaling, spot strategy, and cost efficiency - **Create self-service tooling**: libraries/templates so analytics engineers and data scientists can own pipelines without filing tickets - **Close environment gaps**: staging environment, CI testing against staging, automated schema-change detection, gated promotion, and canary deploys for critical models - **Own governance + access**: Lake Formation row/column RBAC, StrongDM zero-trust access, SSO, audit logging, and PII handling - **Own observability**: Monte Carlo, lineage, alerting, published SLAs; drive incidents to root cause and durable fixes - **Champion infrastructure as code + automation**: Terraform, GitLab CI, GitOps across everything the team runs --- ## What you own vs. our Analytics Engineers You own the **platform**: ingestion, storage, orchestration, compute, access control, observability, and the frameworks on top. Analytics Engineers own the **business-facing layer**: dbt models, golden datasets, metric definitions, and the semantic layer—consuming your platform as a service. --- ## Must-haves - **4+ years** (Senior: **6+**) in data engineering, data platform, or infrastructure engineering - **Strong Python and SQL**, with experience building **frameworks and tooling** others depend on - **Production orchestration experience** (Airflow, Dagster, Prefect) at meaningful scale—including operations, not just DAG authoring - **Hands-on Apache Spark** + distributed-systems fundamentals - **Deep AWS experience** (S3, EKS/ECS, IAM, Glue/Athena or equivalent) - Comfortable with **CI/CD**, **infrastructure as code (Terraform)**, and **GitOps**; familiar with **Kubernetes** and **Docker** - Proven ownership of **reliability**: SLAs, monitoring, alerting, on-call, and post-incident hardening - Daily, hands-on use of **AI coding tools** (Claude Code, Cursor, or equivalent) as a core part of building/operating infrastructure - Strong cross-functional communication to shape data contracts with backend engineering and align with analytics consumers --- ## Nice-to-haves - Production experience with an **open table format** (Apache Iceberg, Delta Lake, Hudi) and lakehouse migration off a classic warehouse - **Streaming experience**:
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.