Principal Software Architect - Data Platform
SecurityScorecard · Hybrid (Austin, TX)
About this role
**Principal Software Architect — Data Platform** **About SecurityScorecard** SecurityScorecard is the global leader in cybersecurity ratings, continuously rating over 12 million companies across 64 countries. Our patented rating technology is used by 25,000+ organizations for self-monitoring, third-party risk management, board reporting, and cyber insurance underwriting—helping organizations find and fix cybersecurity risks across their digital footprint. **About the Role** SecurityScorecard is hiring a **Principal Software Architect** to lead the **system design of our data platform**. Rating millions of companies continuously means ingesting internet-scale measurement data, processing it across streaming, microbatch, and batch paths, storing it so it remains queryable and affordable as it grows, and serving analytics fast enough for real-time exploration. Because the data is the product, **correctness and data quality are critical**—a regression can impact dashboards, scores, and underwriting decisions. This role is an **individual contributor** reporting to the Chief Architect, partnering closely with engineering leadership, Product, and Data Science. **What You’ll Do** - Own end-to-end system design for the data platform (from ingestion through serving) - Define service boundaries and **data contracts** (schema ownership, compatibility rules, handling breaking changes) - Architect the **lakehouse** (table format, partitioning, schema evolution, compaction, metadata growth) - Design an analytical serving layer for three workload classes with conflicting demands: - Low-latency, high-concurrency customer queries - Ad-hoc internal analytics & Data Science exploration - Bulk delivery to external feeds and partners - Set direction on languages and frameworks in the data stack - Engineer data quality and observability into the platform (validation/quarantine, freshness & completeness SLOs, drift detection, automated lineage/metadata) - Design for correctness and reproducibility in the ratings pipeline (backfills, historical restatements when scoring logic changes) - Write **TDRs, design docs, and standards** that set data architecture direction across teams and embed intent into repos - Review TDRs across engineering and provide substantive feedback on architecture and risk - Partner with AI & Front End Architects on data access patterns; mentor senior/staff engineers **Required Qualifications** - 10+ years of software or data engineering experience, including significant time architecting large-scale data platforms - Deep expertise in stream + batch processing at scale (Kafka, Flink, Spark or close equivalents) and strong workload-path judgment - Strong **Python & PySpark**, solid **Java** for Flink stream processing, and enough **Scala** to read/reason about existing Spark codebases - Hands-on production experience designing lakehouse storage (e.g., Parquet; open table formats like Iceberg; partitioning/compaction/schema evolution) - Experience architecting OLAP/analytical serving layers (ClickHouse, Druid, Pinot, BigQuery, Snowflake, or similar) - Strong distributed systems fundamentals (exactly-once vs at-least-once, ordering, backpressure, late/out-of-order data, failure modes) - Track record building data quality, contracts, and observability as engineered system properties (assertions, schema enforcement, lineage in code) - Experience owning large-scale data migrations while preserving history and correctness through cuto
Listing freshness
CronJobs last confirmed this listing 54m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.