Cyber Evaluations Engineer
Anthropic · San Francisco, CA | Washington, DC
About this role
**Cyber Evaluations Engineer** **About Anthropic** Anthropnic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. **About the Role** Build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. Design new evals, run per-release robustness testing, and analyze data on jailbreaks and prompt bypasses. Design probes that detect cyber abuse in production and help shape the overall detection architecture. **Key Responsibilities** • Design and run capability, uplift, and safety evaluations to assess cyber-relevant risk in new models • Execute per-release safeguard-robustness testing ahead of major model launches • Analyze evaluation results and communicate findings to the team and stakeholders • Design, prototype, and tune detection probes for cyber misuse • Work with the cyber policy team to build a layered, robust abuse-detection architecture • Build and maintain internal tooling used to run and score evaluations • Collaborate with policy and engineering partners to translate eval findings into safeguard improvements **Minimum Qualifications** • Experience building or running evaluations, benchmarks, or test suites for software or ML systems • Hands-on cybersecurity experience (CTF participation, vulnerability research, exploit development, or security research) • Proficiency in Python • Strong ability to communicate evaluation results with cross-functional stakeholders **Preferred Qualifications** • Deep offensive-security or security-research experience, including AI security benchmarks • Experience analyzing adversarial or abuse data (jailbreaks, prompt bypasses, intrusion telemetry) • Experience working onsite with government partners on testing or evaluation engagements • Experience with AI/ML evaluation frameworks • Familiarity with coordinated vulnerability disclosure practices • Experience testing pre-release software or models under confidentiality constraints • Experience authoring detection content (Sigma, YARA, Suricata, SIEM rules) or building ML-based abuse detection • Active secret security clearance or higher, or eligibility to obtain one **Compensation & Logistics** 💰 Annual Salary: $300,000 – $405,000 USD 🎓 Minimum Education: Bachelor's degree or equivalent 📍 Hybrid Policy: At least 25% in-office time 🌍 Visa Sponsorship: Available **We encourage applications from all qualified candidates, including those who don't meet every qualification.**
Listing freshness
CronJobs last confirmed this listing 13h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.