Staff Software Engineer - Reporting, Data Platform & Observability
OneTrust · Atlanta, Georgia
About this role
**Staff Software Engineer — Reporting, Data Platform & Observability** **About OneTrust** OneTrust’s mission is to enable innovation through the responsible use of data and AI. The OneTrust AI-Ready Governance Platform unifies regulatory intelligence, automation, and connected governance workflows so businesses can move at the speed of AI while ensuring good governance. **The Challenge** Join the Reporting and Data Platform team as a hands-on individual contributor focused on designing, building, operating, and improving distributed backend services and data-processing platforms. **Your Mission** **Technical Ownership and Delivery** - Own complex features and technical improvements from discovery through production rollout. - Investigate ambiguous problems, identify root causes, evaluate trade-offs, and implement pragmatic solutions. - Review code and technical designs; document decisions and system behavior; partner with product and engineering teams. - Use AI-assisted engineering tools (e.g., Devin, Claude, or similar) to accelerate delivery while maintaining production-quality standards. **Backend and Distributed Systems** - Design and implement production services using Java, Spring Boot, and Maven (APIs, async workflows, report generation, aggregation, export, scheduling). - Build event-driven functionality using Kafka and related messaging patterns; work with caching, relational storage, and service integrations. - Improve performance, scalability, fault tolerance, and resource efficiency (retries, idempotency, caching, backpressure, concurrency, failure recovery). - Diagnose issues across services, queues, databases, and downstream dependencies; modernize capabilities incrementally. **Data Engineering** - Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake (batch + streaming). - Implement schema evolution, checkpoint management, deduplication, replay, late-arriving-data handling, and standardized data-layer patterns. - Optimize Spark joins, partitioning, Delta operations, cluster utilization, and query performance; troubleshoot failed/delayed/inefficient workloads. - Protect tenant boundaries across joins/aggregations/deduplication/Delta operations; implement data-quality controls; monitor freshness/completeness/correctness. - Work securely with Azure storage, identities, secrets, and encryption. **Observability, On-Call, and Operational Excellence** - Improve observability across backend services, event-driven workflows, and data pipelines (metrics, structured logs, traces, business telemetry). - Build and maintain dashboards/monitors/alerts using Datadog and Grafana; apply OpenTelemetry, Prometheus, and Micrometer patterns where appropriate. - Participate in on-call rotation and incident response using PagerDuty, Datadog monitors, or equivalent. - Reduce recurring alerts and operational toil by improving alert quality, eliminating noisy monitors, creating runbooks/diagnostic tools, and applying learnings from blameless incidents. - Improve end-to-end correlation and monitor availability, error rates, latency, ingestion lag, data freshness, throughput, consumer lag, job health, rejected records, checkpoint health, tenant-specific failures, data-quality violations, and Spark resource utilization. **What Success Looks Like** - You require limited direction and can turn ambiguous problems into concrete deliverables. - You own work end-to-end: design → implementati
Listing freshness
CronJobs last confirmed this listing 5h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.