CronJobs

backend jobs

RL Environments Engineer

Bespokelabs · Mountain View

onsiteunknown$250,000–$300,000Posted Aug 27, 2026reinforcement-learningpipelinesautomationtestingCI/CDverificationQAtooling

Apply on the employer site

About this role

**RL Environments Engineer — Bespoke Labs** **About Bespoke Labs** Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents. We’ve curated **Open Thoughts** (https://open-thoughts.ai/), one of the best open reasoning datasets used by multiple frontier labs, and trained SOTA specialized models such as **Bespoke-MiniChart-7B** (https://www.bespokelabs.ai/blog/bespoke-minichart-7b) and **Bespoke-MiniCheck** (https://www.bespokelabs.ai/bespoke-minicheck). We also taught multi-turn tool-calling with reinforcement learning (https://www.bespokelabs.ai/blog/improving-multi-turn-tool-use-with-reinforcement-learning-agents). **Role Summary (Delivery-focused)** This is a delivery role. You’ll build the machinery that turns environment ideas into **hundreds or thousands of validated agentic coding tasks**, and you’ll do it **fast**. You won’t be studying environments in the abstract—you’ll build the pipelines that mass-produce them, design the coding worlds agents train inside, and keep pushing throughput (more environments, higher quality, less manual work per task). We’ll measure you on the **volume and quality** of environments you ship—not papers. --- ## What You’ll Do - **Build environment-generation pipelines**: Own systems that produce RL environments programmatically (templating, automated grading, verification, QA) so the team ships at scale. - **Create complex coding worlds**: Build high-fidelity environments around real codebases, including conventions, dependencies, tooling, and realistic technical debt. - **Scale agentic task creation to thousands**: Move from dozens to **hundreds/thousands** of validated agentic coding tasks with automation doing the heavy lifting. - **Build tools that raise throughput**: Identify bottlenecks in environment production and remove them; create internal tooling/infrastructure to speed up the team. - **Own the full task lifecycle**: Prompt → environment → grader → run frontier models → failure analysis → iteration until tasks are rigorous, fair, and hard to game. - **Defend quality at scale**: Catch reward hacking and grader loopholes; build verification and standards as volume grows. - **Direct coding agents heavily**: Use frontier coding agents to build/validate environments faster and catch subtle failures. --- ## What We’re Looking For - **A record of shipped volume**: You’ve built agentic coding tasks or environments—show what you personally drove and what it cost to produce. - **Experience scaling via automation**: Scaling output through automation rather than adding more people doing manual work. - **Strong software engineering fundamentals**: Fluency in multiple languages and production-grade engineering practices. - **Real production software experience**: Large codebases, build systems, testing, deployment, on-call, and root-cause analysis. - **Adversarial mindset**: You look at a grader and ask how a model would cheat it—then fix it. - **Clear understanding of frontier coding agents**: Know what they can/can’t do and where they cut corners. - **Ownership**: Build, debug, and ship with minimal supervision. --- ## You May Be a Good Fit If You Also Have - Experience with **RL training systems**, post-training, verifiers, or tool-use harnesses - Background in **developer tooling**, CI/CD, sandboxes, or code execution infrastructure - Built large-scale automated test generation, fuzzing harnesses, or benchmark suites (even if

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord