HPC Systems Administrator
Hyperscale and AI Data Center and Cloud Computing
Location: Remote | Department: Cloud Ops
Employment Type: Full-time | Reports To: Senior Infrastructure Operations Manager
DESCRIPTION OF POSITION/ROLE SUMMARY
As an HPC Systems Administrator, you will support the day-to-day operation, maintenance, and troubleshooting of HPC and AI infrastructure across data center and cloud environments.
This role focuses primarily on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management.
You will work alongside senior Systems Administrators, Network Engineering, Data Center Operations, Deployment Engineering, vendors, and other technical teams to troubleshoot infrastructure issues, restore systems to service, perform routine maintenance, and improve the reliability of production HPC environments.
This position is well suited for someone with a foundation in Linux systems administration or data center infrastructure who wants to develop deeper expertise in HPC, GPU computing, high-performance networking, and large-scale AI infrastructure.
KEY RESPONSIBILITIES (What You'll Be Contributing)
HPC Infrastructure Operations
- Support and maintain Linux-based HPC and AI compute environments.
- Assist with the operation of large-scale GPU clusters built on NVIDIA HGX, DGX, or similar accelerated computing platforms.
- Troubleshoot common hardware, operating system, driver, firmware, networking, and storage issues.
- Investigate degraded, unavailable, or unstable compute and GPU nodes using established troubleshooting procedures.
- Escalate complex or unfamiliar issues to senior administrators or engineering teams when appropriate.
- Assist with identifying recurring infrastructure issues and documenting technical findings.
- Support bare-metal, virtualized, containerized, and cloud-hosted systems.
- Maintain operational documentation, support procedures, runbooks, and troubleshooting guides.
GPU and Accelerated Computing Systems
- Assist with the installation, configuration, validation, and troubleshooting of NVIDIA GPU drivers, firmware, and supporting system software.
- Monitor and troubleshoot GPU health using tools such as nvidia-smi, DCGM, and NVIDIA Fabric Manager.
- Identify and collect diagnostic information for GPU Xid errors, NVLink issues, PCIe errors, ECC events, GPU resets, and thermal or power-related issues.
- Assist senior team members with troubleshooting multi-GPU communication and performance issues.
- Perform post-maintenance validation of GPU systems following repairs or infrastructure changes.
- Coordinate and track hardware replacement and RMA activities for GPUs, system boards, NICs, power supplies, and other server components.
Linux Systems Administration
- Administer and support enterprise Linux systems.
- Troubleshoot common boot, service, filesystem, memory, CPU, device-discovery, and operating-system issues.
- Support Linux networking, storage mounts, authentication, permissions, SSH, DNS, and NTP.
- Collect diagnostic information for system crashes, kernel issues, hardware errors, and out-of-memory conditions.
- Follow established security and operating-system configuration standards.
Server Hardware, Firmware, and Out-of-Band Management
- Support enterprise server platforms from vendors such as Dell, HPE, ASUS, Supermicro, NVIDIA, or similar manufacturers.
- Troubleshoot common hardware issues involving memory, PCIe devices, GPUs, NICs, storage devices, power supplies, fans, system boards, and cabling.
- Assist with BIOS, BMC, NIC, GPU, drive, and other component firmware updates.
- Use out-of-band management tools to perform remote console access, power operations, hardware inventory collection, log collection, and system health checks.
- Work with Data Center technicians during hardware replacements, cabling validation, break-fix activities, and post-repair testing.
- Follow established procedures for firmware upgrades and hardware maintenance.
High-Performance Networking
- Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments.
- Support NVIDIA/Mellanox and other network adapters used in compute infrastructure.
- Perform basic troubleshooting of link state, interface errors, MTU configuration, packet loss, VLANs, RDMA, RoCE, and InfiniBand connectivity.
- Collect network diagnostic information and escalate complex fabric or switching issues to Network Engineering or senior team members.
- Assist with validation following network changes, firmware upgrades, cable replacements, and new deployments.
Incident, Problem, and Change Management
- Participate in incident response for production infrastructure events.
- Assist with troubleshooting activities and provide clear technical updates during incidents.
- Document troubleshooting steps, findings, actions taken, and resolution details.
- Escalate incidents according to established support and escalation procedures.
- Follow defined change-management processes, maintenance procedures, and rollback plans.
- Work within established service-level agreements and operational processes.
- Participate in a scheduled on-call rotation as required.
Collaboration and Development
- Work closely with senior Systems Administrators and engineers to develop HPC infrastructure troubleshooting skills.
- Follow established troubleshooting methods, documentation standards, and operational procedures.
- Contribute to runbooks, knowledge-base articles, and support documentation.
- Participate in technical reviews, troubleshooting sessions, and post-incident discussions.
- Share technical findings and lessons learned with other team members.
- Continue developing knowledge of Linux, GPU infrastructure, high-performance networking, storage, and automation.
QUALIFICATIONS REQUIRED (What Sets You Apart)
- Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role. Equivalent hands-on experience, education, or technical training may also be considered.
- Working knowledge of Linux operating systems and command-line administration.
- Basic understanding of enterprise server hardware and common server components.
- Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools.
- Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting.
- Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware.
- Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms.
- Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools.
- Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies.
- Ability to follow technical procedures and troubleshoot problems methodically.
- Ability to recognize when an issue requires escalation and clearly communicate collected findings.
- Strong written communication and documentation skills.
- Ability to work collaboratively across technical teams.
- Participate in an on-call rotation.
STRONGLY PREFERRED
- Exposure to NVIDIA HGX, DGX, H100, H200, B100, B200, or similar GPU platforms.
- Familiarity with nvidia-smi, DCGM, NVIDIA Fabric Manager, CUDA, or GPU diagnostic tools.
- Exposure to InfiniBand, RDMA, RoCE, NVIDIA/Mellanox adapters, or NVIDIA UFM.
- Familiarity with NVLink, NVSwitch, NCCL, or multi-GPU systems.
- Experience working with enterprise server hardware from Dell, HPE, Lenovo, Supermicro, NVIDIA, or similar vendors.
- Exposure to MAAS or other bare-metal provisioning technologies.
- Familiarity with NFS, Lustre, BeeGFS, Spectrum Scale, Ceph, or other storage technologies.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, DCGM Exporter, or similar platforms.
- Experience with hardware diagnostics, vendor support cases, or RMA processes.
- Familiarity with Git, Ansible, Terraform, or other infrastructure automation tools.
- Familiarity with Jira Service Management or another ITSM platform.
- Certifications such as RHCSA, CompTIA Linux+, Linux Foundation, NVIDIA, CCNA, or equivalent technical training.
WORKING CONDITIONS
- Remote-first position with extensive computer-based work.
- Work may involve coordinating remote troubleshooting and hands-on activities with Data Center technicians.
- Fast-paced, customer-facing production environment supporting business-critical HPC and AI infrastructure.
- Participation in a scheduled on-call and escalation rotation is required.
- Occasional work during maintenance windows, weekends, or outside normal business hours may be required.
- Additional responsibilities may be assigned based on operational requirements and company growth.
5C Data Centers is an equal opportunity employer.
At 5C Data Centers, we celebrate diversity and are committed to creating an inclusive environment where everyone can thrive.
5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.
#LI-AJ1