Site Reliability Engineer
Accela · Remote Based - US
About this role
<p><strong>ABOUT THE ROLE:&nbsp;</strong></p> <p>Accela provides cutting-edge technology that enables government agencies to engage and serve their communities. The foundation of our technology is the Accela Civic Platform, a cloud-based platform that supports a broad ecosystem of solutions for permitting, licensing, planning, and public sector operations.</p> <p>As a Site Reliability Engineer (SRE), you will be responsible for maintaining, monitoring, and improving the reliability, availability, performance, and scalability of Accela's cloud platform and SaaS services. You will work across cloud infrastructure, applications, and operational processes to ensure a seamless customer experience through proactive monitoring, incident response, automation, and continuous improvement.</p> <p>This role combines cloud operations, production support, and reliability engineering, with a strong emphasis on hands-on administration of Microsoft Azure services, observability platforms, incident management, and operational excellence.</p> <p><strong>SPECIFIC RESPONSIBILITIES:&nbsp;</strong></p> <p>Reliability Engineering &amp; Platform Operations</p> <ul> <li>Monitor and maintain production cloud environments to ensure platform availability, performance, scalability, and reliability.</li> <li>Build, configure, and optimize monitoring, logging, and alerting capabilities using Datadog.</li> <li>Develop and maintain dashboards that provide real-time visibility into platform health, resource utilization, and service performance.</li> <li>Configure and tune alerts to proactively identify and address service degradation or operational issues.</li> <li>Support and optimize Azure-based infrastructure components, including Azure Kubernetes Service (AKS), Azure SQL, Azure Storage, and Azure Front Door.</li> <li>Execute production releases, operational changes, and platform maintenance activities.</li> </ul> <p>Incident Management &amp; Continuous Improvement</p> <ul> <li>Participate in and lead incident response activities to restore service and minimize customer impact.</li> <li>Diagnose and resolve production incidents across infrastructure, platform, and application components.</li> <li>Conduct root cause analysis (RCA) and implement corrective and preventative actions to reduce recurring issues.</li> <li>Develop, maintain, and continuously improve operational documentation, runbooks, and incident response procedures.</li> <li>Identify and implement automation opportunities that improve operational efficiency and system reliability.</li> </ul> <p>Customer Support &amp; Operational Excellence</p> <ul> <li>Provide Level 3 (L3) support for customer-reported incidents and service requests.</li> <li>Collaborate with Engineering, Professional Services, and Operations teams to investigate and resolve complex technical issues.</li> <li>Meet established service level agreements (SLAs) for ticket response and resolution.</li> <li>Support customer provisioning, environment maintenance, platform upgrades, and cloud migration activities as needed.</li> </ul> <p>Data &amp; Environment Operations</p>...
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.