Incident Operations Commander
Alpaca · Remote - Americas
About this role
**Who We Are** Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and 24/5 trading. We serve hundreds of financial institutions across 40 countries with institutional-grade APIs—supporting broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges (10M+ brokerage accounts). **Role: Incident Operations Commander** Serve as the on-duty commander for Alpaca’s most critical incidents—directing cross-functional response to restore service quickly, keeping the right people engaged and informed, and ensuring every incident produces actionable follow-up. > You do not fix the outage. You make the response reliable: correct severity, the right engineers in the room, mitigation that doesn’t stall, leaders informed in time, and follow-up work that survives the call. **Things You Get To Do** - **Command incidents end to end:** Take command from declaration to mitigation; run the bridge; keep observers out of the responders’ way; call out stalls. - **Classify and hold the line on severity:** Set severity at declaration and re-check as facts arrive (risk advises on materiality—the call is yours). - **Engage the right people, fast:** Identify the owning team by service/symptoms/blast radius; page and expand the responder set as needed; escalate when pages go unanswered. - **Hold the bridge and protect the people fixing it:** Be the single source of truth for partner communications on impact/severity/timing; decide when updates are needed. - **Run follow-the-sun handoffs:** Deliver warm, high-fidelity handoffs across regions; ensure command never goes dark at region boundaries. - **Close the loop, on the clock:** Maintain the incident timeline; schedule a blameless retrospective with an owner and timebox; ensure follow-ups become real tickets with accountable owners, priority, and SLA delivery. - **Automate coordination away:** Use decision trees and agentic automation to reduce manual toil so commanders can focus on judgment. **Who You Are (Must-Haves)** - **4+ years** commanding/co-commanding high-severity incidents in production engineering, SRE, or technical operations. - You can **direct technical responders under pressure** without being the person writing the fix. - You can **make and defend crisp severity and escalation decisions** and take charge immediately (including waking senior leaders at 03:00 and directing experienced engineers to stop). - You can **read dashboards** and judge whether impact has truly stopped. - You communicate clearly with **engineers, executives, and partner-facing stakeholders**, and know the difference between briefing comms vs. speaking for the company. - You can hold other teams accountable in the moment across reporting lines without creating friction. - You thrive in a **follow-the-sun** model with clean cross-region handoffs. - You understand **FinTech concepts** and the trust stakes of API-driven financial platforms. - You use **AI tools and agentic automation** to reduce manual toil and speed response. - You will work a **regional coverage window** as part of a global **24/7** Incident Commander roster. **Who You Might Be (Nice-to-Haves)** - Formal incident command training (ITIL, Major Incident Management, crisis management). - Experience with modern incident management and on-call platforms. - Experience writing **severity rubrics, decision trees, escalation matrices, runbooks, or incident play
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.