Cloud Systems Engineer
Alarm.com · Tysons, Virginia
About this role
**Cloud Systems Engineer** We are seeking a **Cloud Systems Engineer** to support and operate large-scale **AI** and **high-performance computing (HPC)** environments. This role focuses on the deployment, maintenance, performance, and lifecycle management of **GPU-accelerated compute infrastructure** powering critical AI, machine learning, and data-intensive workloads. --- ## Responsibilities ### AI Infrastructure Operations - Deploy, configure, and maintain **GPU-accelerated compute infrastructure** - Manage OS, firmware, BIOS, BMC, driver, and software lifecycle updates - Monitor system health, performance, utilization, and capacity across AI infrastructure environments - Support infrastructure used for **AI model training, inference, and data processing** - Develop and maintain operational standards, **runbooks**, and maintenance procedures - Participate in **on-call support** and incident response ### Linux Systems Administration - Administer enterprise Linux environments (including **Ubuntu** and **Red Hat-based** distributions) - Perform system patching, hardening, and OS lifecycle management - Troubleshoot OS, kernel, storage, networking, and application-level issues - Develop automation to streamline deployment, monitoring, and operational processes - Support security and compliance initiatives across AI infrastructure platforms ### Hardware and Datacenter Operations - Install, configure, maintain, and troubleshoot enterprise compute hardware - Diagnose and resolve issues involving **GPUs, CPUs, memory, storage, power, and networking** - Perform firmware upgrades and hardware lifecycle management - Coordinate hardware replacements, vendor support engagements, and warranty services - Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects - Maintain accurate asset inventories and operational documentation - Support high-performance networking (including **Ethernet** and **InfiniBand**) - Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations - Assist with scalability, resiliency, and performance optimization initiatives - Perform root-cause analysis of infrastructure failures and develop preventative measures --- ## Required Qualifications - **Bachelor’s degree required** - **3–5 years** of Linux systems administration experience in production environments - **3–5 years** supporting enterprise server infrastructure - Experience supporting large-scale compute environments, **HPC**, AI infrastructure, or **GPU-enabled systems** - Experience performing hardware diagnostics, firmware management, and lifecycle maintenance - Experience working within datacenter operations environments - Scripting: **Bash, Python, PowerShell**, or similar - OS performance tuning and monitoring - Storage and networking fundamentals - Experience with infrastructure monitoring/observability platforms (e.g., **Grafana**) - Hardware and firmware lifecycle management --- ## Sponsorship Notice Please note that sponsorship of new applicants for employment authorization, or any other immigration-related support, is **not available** for this position at this time. --- ## Why Work for Alarm.com? - Collaborate with outstanding people and high standards - Make an immediate impact with real responsibility - Gain well-rounded experience across multiple areas of the business - Community and camaraderie (“Keep It Fun”) - In-person collaboration: employees work f
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.