STAFF SOFTWARE ENGINEER, Frontier Security Team
Snowflake · US-CA-Menlo Park
About this role
## Staff Software Engineer — Frontier Security Team At Snowflake, we’re powering the era of the **agentic enterprise**. We’re seeking **AI-native thinkers** who are energized by reinventing how work gets done—using AI as a high-trust collaborator, moving with an experimental mindset, and delivering measurable results. Snowflake’s **Frontier Security AI teams** develop **production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems** for enterprise customers. Our products must meet high standards for **quality, security, reliability, and efficiency** while operating over **sensitive data at large scale**. --- ## About the Role We’re looking for a **Staff Software Engineer** to lead the design and development of our **Agentic Harness** and **agent evaluation platform**. The **Agentic Harness** provides the runtime, tools, context, state, policies, and observability required to build and operate production AI agents. The **evaluation platform** measures how well agents complete **real customer tasks** and detects regressions **before** they reach production. This is a **hands-on technical leadership** role where you’ll build production systems, establish architecture across team boundaries, and define the metrics and engineering practices used to improve agent quality. --- ## What You’ll Do - **Architect and build** the Agentic Harness to execute complex, multi-step AI workflows across models, tools, data, and services. - Design stable interfaces for **tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review**. - Own **agent quality end to end** by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. - Turn ambiguous reports (e.g., “the agent feels worse”) into **measurable failure modes**, **reproducible tests**, and **durable fixes**. - Analyze production agent trajectories to identify failures in **reasoning, retrieval, tool use, context, orchestration, and application code**. - Close the loop between **production incidents**, **root-cause analysis**, **evaluation coverage**, and **regression prevention**. - Develop **offline and online measurements** for task completion, correctness, groundedness, safety, latency, reliability, and cost. - Build **simulation and replay infrastructure** for golden-set tests, adversarial scenarios, model comparisons, and large-scale experiments. - Improve agent efficiency via **model routing, prompt/semantic caching, context compaction, tool-result management, and token optimization**. - Productionize new model capabilities as **secure, observable, multi-tenant services** with clear operational controls. - Establish standards for evaluation design: **sampling, ground-truth quality, grader calibration, leakage prevention, and statistical significance**. - Define technical direction across multiple teams; lead projects beyond a single service. - Mentor engineers, raise the quality of architecture reviews, and stay directly involved in implementation and debugging. --- ## Requirements - **9+ years** of software engineering experience, including technical leadership of complex production systems. - Experience shipping and operating **LLM applications, AI agents, or model-backed workflows** in production. - Strong background in **distributed systems**, service architecture, high-throughput APIs, concurrency, and failure handling. - Experience buil
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.