Sr. Staff Machine Learning Systems Engineer
hims & hers · USA
About this role
## Sr. Staff Machine Learning Systems Engineer ### About the Role Hims & Hers is looking for a **Senior Staff ML Systems Engineer** to make advanced AI/ML **powerful and trustworthy** for a **regulated healthcare** environment. You’ll own the end-to-end evaluation and data infrastructure that determines whether AI models are **effective and safe to ship to patients**. This is a leadership role spanning disciplines—setting technical direction for how we **build, version, and trust** the data and judgments used to evaluate AI products. Your scope runs from **data ingestion and feature/dataset pipelines** through **evaluation methodology and statistical rigor**, to the **reporting surfaces** that help clinical and product teams act on what we learn. ### You Will - **Own evaluation end-to-end** (not just a slice) - Set technical direction for evaluation systems, including: - **Metric, judge, and scorer design** - **Statistical methodology** behind regression decisions - Infrastructure to **track and categorize failures over time** - **Design and scale data pipelines** for evaluation and downstream work, including: - Ingestion, transformation - Dataset versioning - Labeling and calibration workflows - Proactively address challenges in **scaling and complexity** of AI evaluation - Lead projects spanning teams and quarters, including defining evaluation for new AI services **from scratch** - Drive multi-team initiatives (e.g., replacing **manual, inconsistent review** with **statistically sound automated gates**) - Own adversarial and **red-team evaluation** as a risk-reduction program: - Design test suites and failure taxonomies to catch safety/edge-case issues before patients - Resolve complex cross-team technical disagreements and drive alignment across engineering, product, and AI leaders - Turn ambiguous, cross-team pain points into **fully specified, production-ready systems** - Lead major platform improvements (re-architecting core systems, removing brittle logic, modernizing operations) - Share knowledge through talks and write-ups; help shape reusable standards - Mentor engineers and raise the technical/statistical bar across teams ### You Have - **10+ years** experience in ML infrastructure, data engineering, or evaluation/testing systems, with impact beyond a single team - Hands-on depth in: - Evaluation systems (e.g., **designing/calibrating LLM judges/scorers**) - Statistically sound regression-testing methodology (e.g., paired significance testing with **multiple-comparisons correction**) - Measuring agreement against human labels - Adversarial/red-team evaluation approaches - Hands-on depth in data pipeline engineering, including: - Dataset versioning - Feature/benchmark pipelines - Labeling and calibration workflows - High-throughput ingestion and transformation systems - Experience building approaches that become **standards** for others - Experience leading multi-team projects to completion, including navigating technical disagreement - Track record mentoring engineers (including senior engineers) and raising team standards - Strong communication skills across audiences - Strong **Python** and statistical fluency to design and defend testing frameworks used for production decisions ### Nice to Have - Experience with **Databricks, MLflow, Unity Catalog**, or similar platforms - Experience building reporting tools for non-engineering stakeholders (e.g., automated Slack digests, spreadsheet rep
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.