Software Engineer, Applied AI
Claylabs · New York
About this role
## About Clay AI Clay is unleashing the biggest wave of company creation in history. Our mission is to be the engine companies use to grow to their full potential—predicting the next best action and helping teams take it. We started by aggregating the best data for B2B companies, then built infrastructure to run personalized campaigns (emails, ads, landing pages, and workflows that make reps more productive). Now we’re building **AI agents** that help grow your company. Clay already supports thousands of customers, including **Anthropic, OpenAI, Google, and Visa**. ## About the Team Clay’s product is increasingly powered by AI agents—systems that research, enrich, and take action on behalf of users (not just generate text). This role is a shared entry point across teams working on: - **Agent products** that execute real go-to-market workflows end-to-end - A **shared agent platform** (harness, memory, tools, retrieval, evals) All teams are focused on one underlying challenge: closing the gap between an agent that looks good in a demo and one that’s dependable enough to run unattended in production. ## About the Role You’ll work closely with product and research-adjacent teammates, plus other engineers, to ensure agents are **reliable, steerable, and worth trusting** with real work. This isn’t only about improving model behavior—it’s about turning improvements into measurable gains in: - task completion - reliability - time saved for users ## What You’ll Do Depending on the team, you may work on: ### Agent Products - Design and iterate on agent behavior across real GTM workflows (e.g., sourcing a TAM list by navigating ambiguous ICP definitions, reconciling conflicting signals, and making judgment calls) - Map manual, multi-step GTM workflows into agent-driven flows that are as good as—or better than—human execution - Build and run **evals** to measure whether an agent completed tasks correctly (not just whether outputs look plausible), and use them to catch regressions and failure modes - Analyze real production failures and improve robustness - Partner with product to move agent flows from prototype → closed beta → general availability, and help define what “good” means ### Agent Platform & Infrastructure - Build the core agent harness (memory systems, tool infrastructure, retrieval architecture) - Improve agent performance via prompting strategies, tool-use design, and context construction - Design guardrails and safety checks so agents behave predictably in production - Create a cross-surface evals framework so every team can measure quality/regressions/performance consistently - Build feedback loops that turn real usage and production logs into better prompts, tools, and eval coverage over time - Support teams building forks/variants of managed agents for specific use cases ## What You’ll Bring - Experience building/shipping **production systems** with LLMs or agents (multi-step orchestration, prompting + tool-use design, retrieval, structured extraction, or fine-tuning) - Strong backend fundamentals: APIs, databases, distributed systems - Experience with **model/agent evaluation** (designing evals, measuring regressions, turning fuzzy quality into measurable signals) - A systems-and-outcomes mindset (care whether the product works for users) - Comfort debugging messy, real-world failures and a bias toward shipping + iterating quickly ## Nice to Haves - Experience with agent frameworks, tool-calling systems, or retrieval
Listing freshness
CronJobs last confirmed this listing 14h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.