AI Data Engineer, Data Platform
Collective · San Francisco
About this role
## About Collective Collective is on a mission to redefine how businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business incorporation to accounting, bookkeeping, tax services, and access to a thriving community—all in one integrated platform. We believe self-employed people should enjoy the same tax savings as big companies, so they can focus on their passion—not paperwork. Featured in Forbes, Business Insider, Yahoo, Bloomberg, Financial Times, TechCrunch, and more. Backed by General Catalyst, Sound Ventures, QED Investors, Google’s Gradient Ventures, Expa, and other investors. --- ## About the Role (AI Data Engineer, Data Platform) We’re looking for a **Data Engineer** to own and scale the data platform that powers **analytics, reporting, and AI** across Collective. You’ll design, build, and maintain pipelines that move data from product, financial, and third-party systems into **BigQuery**; model that data into clean, well-documented, reliable tables; and set engineering standards that keep the platform trustworthy as the company grows. This is a hands-on role within the **Data Engineering team** (Engineering), working closely with product engineers, analysts, and stakeholders across Operations, Finance, and Go-to-Market. --- ## What You’ll Do - **Design & build data pipelines**: scalable batch and event-driven pipelines ingesting data into BigQuery using managed connectors (e.g., **Fivetran**), custom **Python** loaders, and orchestration tooling. - **Model the data**: build dimensional and analytical models in **dbt** using a layered architecture (**raw → staging → marts**) with clear grain, naming conventions, and documentation. - **Own data quality & reliability**: implement testing, monitoring, alerting, and data contracts; define and meet **freshness/accuracy SLAs**; triage and resolve incidents to root cause. - **Optimize performance & cost**: tune queries, partitioning, and clustering; manage BigQuery spend; keep pipelines efficient as data grows. - **Establish engineering standards**: drive best practices for version control, code review, CI/CD, and infrastructure-as-code; document systems and runbooks. - **Govern & secure data**: implement access controls, **PII handling**, and data retention practices appropriate for a financial services company; partner with Security and Legal. - **Enable the business**: partner with product engineers on schema design and change management; work with analysts/stakeholders to translate questions into reliable datasets, metric definitions, and self-serve reporting in **Metabase**. - **Support AI & analytics use cases**: maintain the semantic layer, metric definitions, and documentation so LLM-based tools and internal agents can query accurately and consistently. --- ## What You’ll Bring - **Experience**: 5+ years in data engineering / analytics engineering (ideally B2B SaaS or fintech). - **SQL & Python**: expert-level SQL and strong Python for tested, production-grade pipelines and tooling. - **Modern data stack**: production experience with **BigQuery** (preferred), **dbt**, managed ingestion (e.g., Fivetran), and an orchestrator (Airflow/Dagster/Cloud Composer/etc.). - **Data modeling**: strong dimensional modeling and layered warehouse architecture; clear opinions on grain, naming, and consistency. - **Data quality & observability**: testing frameworks, lineage, monitoring, alerti
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.