CronJobs

devops-sre jobs

Major Incident Management Specialist

Peraton · San Antonio, TX, US

onsiteunknown$86,000–$138,000Posted Sep 13, 2026ServiceNowITILITSMActive DirectorySaaSIaaSVPNCommVault

Apply on the employer site

About this role

## Major Incident Management Specialist ### Responsibilities - Provide coordinated support for major incident process components, including outage reporting, documentation, training, investigation, and quality control. - Work within the DHA Global Network Operations Center (GNOC) to support rapid troubleshooting for outages impacting Military Treatment Facilities (MTFs) supported by the Defense Health Agency (DHA). - Collaborate with cross-functional teams (Event Management, Problem Management, Incident Management, Service Reporting, DHA Infrastructure & Operations, JOMIS, DHMSM, DISA, DMDC, NIWC, Coast Guard, VA, PMO application owners, and vendors) to restore IT services impacting healthcare delivery. - Use Government-provided ITSM tools (currently ServiceNow), knowledge bases, and ITIL-based processes to coordinate service owners and drive resolution. - Facilitate service restoration: contact service owners, document troubleshooting steps and root causes, create downtime notifications, and interface with GNOC Government Watch Officers. - Accountable for the Major Incident Management Process and alignment to key service partners. - Facilitate and complete actions within the Continual Service Improvement Plan for Major Incident Management. - Ensure end-to-end ownership and accountability for each major incident. - Manage communications for major incidents: define proactive communications plans, review throughout, and ensure conference calls occur when appropriate. - Track a defined timeline during major incident events. - Maintain close links with Problem Management; ensure problems are raised appropriately after major incidents. - Identify workarounds and known errors to support Problem Management. - Partner with Event Management (monitoring) teams to understand event alerts and establish service restoration response. - Promote and educate internal/external resolution teams on the Major Incident Management process and outage reporting steps (e.g., AOT’s). - Support post-incident review input for critical incidents related to outage reporting. - Collaborate with CommCell and GNOB leadership to maintain consistent operational visibility and alignment. - Ensure outage reports, incident details, timelines, and restoration actions are documented clearly and accurately. - Support post-incident review activities by gathering incident data, technical notes, service owner feedback, and RCA inputs. - Maintain current contact information, service catalogs, and critical business service references. ### Skills / Qualifications - Experience managing critical incidents and problem records in a call center environment. - Ability to create and deliver executive summary and after-action reports. - Experience training groups at multiple organizational levels. - Strong interpersonal skills; ability to interface with internal/external customers, vendors, and management. - Excellent verbal and written communication. - Knowledge of ITIL 4/5 and experience with ITSM solutions. - Demonstrated problem coordination and root cause determination. - Ability to understand system/solution architectures to identify Configuration Items (CIs) and impacted services. - Technical familiarity in one or more domains (examples): - Local MTF Networks - Cybersecurity Tools - Network Circuits & VPNs - Active Directory - SaaS/IaaS - End-User Services - CommVault / SCCM - JOMIS/MHS Medical Application Architecture and Support ### Required Qualifications - Bache

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord