CronJobs

backend jobs

AI Evaluation Infrastructure Engineer

Block (Square) · Bay Area, CA, United States of America

remoteunknown$263,600–$263,600Posted Sep 30, 2026

Apply on the employer site

About this role

<p>It all started with an idea at Block in 2013. Initially built to take the pain out of peer-to-peer payments, Cash App has gone from a simple product with a single purpose to a dynamic ecosystem, developing unique financial products, including Afterpay/Clearpay, to provide a better way to send, spend, invest, borrow and save to our 50+ million monthly active customers. We want to redefine the world’s relationship with money to make it more relatable, instantly available, and universally accessible.<br><br>Today, Cash App has thousands of employees working globally across office and remote locations, with a culture geared toward innovation, collaboration and impact. We’ve been a distributed team since day one, and many of our roles can be done remotely from the countries where Cash App operates. No matter the location, we tailor our experience to ensure our employees are creative, productive, and happy.</p> <h4><strong>The Role</strong></h4> <p>We build AI products, and the quality of our evaluations sets the ceiling for how good those products can be. The speed of our evaluations determines how quickly we can improve them.</p> <p>We are looking for an engineer to build the infrastructure and tooling that make high-quality AI evaluation possible at Block's scale. You will help teams understand whether a model or product change is actually better, whether a result is statistically meaningful, and whether offline evaluation is predicting what happens with real users.</p> <p>Our evaluation approach combines offline evals that encode our definition of a good response, online evals that show how people actually respond, and a feedback loop that keeps the two converging. Your work will turn that approach into systems that product teams can use quickly, reliably, and with confidence.</p> <p>This is a high-impact, early-stage area with broad surface area. You will help decide what to build first, then build the platform that helps teams ship better AI products faster.</p> <h4><strong>You Will</strong></h4> <ul> <li>Build an execution engine that can score candidate versions against task sets in minutes, not hours.</li> <li>Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.</li> <li>Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.</li> <li>Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.</li> <li>Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance, so teams can distinguish real improvements from noise.</li> <li>Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.</li> <li>Build the loop that compares offline scores with online outcomes, identifies eval sets that have stopped predicting reality, and helps teams improve them.</li> <li>Partner with product, engineering, data, and ML teams to make evaluation workflows fast enough and trustworthy enough to become part of everyday development.</li> </ul> <h4><strong>You Have</strong></h4> <ul> <li>Experience building...

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord