CronJobs

backend jobs

Senior/Staff FDE - Synthetic Data Generation

Snorkel AI · New York City, NY (Hybrid); San Francisco, CA (Hybrid)

hybridsenior$180,000–$320,000Posted Aug 27, 2026PythonLLMDockerAWSGCPAzureMLGenAI

Apply on the employer site

About this role

**Senior/Staff FDE - Synthetic Data Generation** **About Snorkel** At Snorkel, we believe meaningful AI starts with the data. We're on a mission to help enterprises transform expert knowledge into specialized AI at scale. **About the Role** Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives. You will lead technical execution of complex customer engagements, translate model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to improve data quality and model outcomes. **Main Responsibilities** **Synthetic Data Generation & Evaluation** • Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines • Translate model objectives and data gaps into synthetic data strategies and technical specifications • Develop LLM- and ML-assisted workflows for high-quality training and evaluation datasets • Build automated evaluators and quality checks to assess correctness, relevance, diversity, and coverage • Design and run experiments to measure synthetic data impact on model performance • Package and deliver production-grade datasets with standardized formats and documentation **Forward Deployed Engineering & Customer Partnership** • Lead technical workstreams from solution design through production delivery • Build and iterate on solutions that address customer needs • Rapidly prototype and productionize solutions across models, pipelines, and APIs • Communicate technical tradeoffs and recommendations to stakeholders • Serve as a trusted technical partner to customers and delivery teams **Technical Leadership & Scale** • Identify patterns across engagements and turn solutions into reusable pipelines and best practices • Define technical standards for synthetic data generation and evaluation • Partner with engineering and product teams to influence platform capabilities • Lead technical design reviews and provide guidance to other engineers • Stay current with emerging synthetic data and LLM evaluation techniques **What We're Looking For** • 5+ years in machine learning engineering, data science, applied AI, or similar technical role • Strong Python skills and experience building production data/ML systems with Docker and cloud platforms • Hands-on experience with LLMs and modern GenAI/LLM stack • Strong understanding of ML experimentation and evaluation • Experience building synthetic data, data augmentation, or model-generated datasets • Experience with LLM evaluation techniques (LLM-as-a-judge, rubric-based evaluation, custom evaluators) • Ability to take ambiguous problems from definition through delivery with strong technical communication • Experience serving as a technical lead with mentoring and architecture influence **Preferred Qualifications** • Experience developing datasets for fine-tuning, preference optimization, or benchmarking • Experience building agentic environments and tasks • Experience with reinforcement learning for LLMs • Experience in fast-paced, customer-facing environments **Compensation** Base salary: **$180,000–$320,000 USD** + variable compensation, equity, and benefits. Actual compensation based on skills, experience, location, and other factors.

Listing freshness

CronJobs last confirmed this listing 11h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord