Software Engineer II - Recommendations
Klaviyo · Boston, MA
About this role
**Software Engineer II – Recommendations** **Location:** Boston, MA (onsite 5x a week) --- ## Why you should join the Recommendations Platform Team The Recommendations Platform Team builds and deploys machine learning-based recommendation systems at scale. You’ll help evolve the platform into a layer of product intelligence—turning customer data into personalized experiences across multiple channels (e.g., email, SMS, KAgent, onsite). The team works across large-scale data pipelines, batch training and inference, and low-latency online retrieval and ranking. You’ll also contribute to experimentation, tracking, and measurement to evaluate recommendation quality and business impact over time. --- ## How you will make a difference - Contribute to the architecture and evolution of **backend services** powering product recommendations across Klaviyo experiences, meeting standards for reliability, performance, and clear APIs. - Contribute to and maintain **robust, large-scale data processing pipelines** (e.g., Apache Spark or similar) that transform events and catalog data into high-quality features for recommendation models, ensuring data quality and lineage. - Collaborate with **ML engineers and product stakeholders** to productionize recommendation models—defining interfaces, feature contracts, and deployment patterns for batch and/or real-time inference. - Contribute to the development of the **vector database** powering recommendations, semantic search, and agentic use cases. - Ensure **data and service observability** (metrics, logging, tracing, dashboards) so recommendations are correct, explainable, fast, and highly available. - Work with Product to break projects into clear milestones—balancing rapid experimentation with technical soundness and long-term maintainability. - Lead **data-driven decision making and A/B testing**—ensuring systems are instrumented with the right metrics and interpreting results to guide future iterations. - Participate in **on-call and incident response** for systems you own, driving post-incident improvements to resilience and operability. - Integrate AI into your and the team’s development workflow (e.g., accelerating development, automating complex tests, improving monitoring/debugging). - Share knowledge, mentor junior engineers, and define best practices for large-scale data frameworks, distributed systems, and integrating ML into production. --- ## Who you are - **2+ years** of professional software engineering experience focused on backend and distributed systems at scale; proven experience building production services and optimizing for latency, reliability, and operability. - Proficient in **Python** (open to working in other languages). - Comfortable with **cloud-native architectures** (AWS preferred) and **container orchestration** (e.g., Kubernetes); able to manage infrastructure and CI/CD as part of development. - Experience with **data-driven decision making and A/B testing**—able to define (or learn) how to instrument experiments, interpret results, and feed learnings back into system design. - Comfortable designing and querying data models in **relational, analytical, and NoSQL** datastores (e.g., Postgres, MySQL, data warehouses, Redis, vector databases). - Familiar with **modern DevOps practices** (CI/CD, monitoring, alerting) and applying them to large-scale data and recommendation systems. - Track record of **owning features end-to-end** (design → implementation → rollout → monit
Listing freshness
CronJobs last confirmed this listing 15h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.