For Employers
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Support and maintain Linux-based HPC and AI infrastructure, including GPU servers, networking, storage, firmware, and out-of-band management across data center and cloud environments. Troubleshoot and document incidents, assist with maintenance and hardware replacements, validate systems after changes, and collaborate with engineering and data center teams.

HPC Systems Administrator

Hyperscale and AI Data Center and Cloud Computing

Location: Remote | Department: Cloud Ops

Employment Type: Full-time | Reports To: Senior Infrastructure Operations Manager

DESCRIPTION OF POSITION/ROLE SUMMARY

As an HPC Systems Administrator, you will support the day-to-day operation, maintenance, and troubleshooting of HPC and AI infrastructure across data center and cloud environments.

This role focuses primarily on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management.

You will work alongside senior Systems Administrators, Network Engineering, Data Center Operations, Deployment Engineering, vendors, and other technical teams to troubleshoot infrastructure issues, restore systems to service, perform routine maintenance, and improve the reliability of production HPC environments.

This position is well suited for someone with a foundation in Linux systems administration or data center infrastructure who wants to develop deeper expertise in HPC, GPU computing, high-performance networking, and large-scale AI infrastructure.

 

KEY RESPONSIBILITIES (What You'll Be Contributing)

HPC Infrastructure Operations

  • Support and maintain Linux-based HPC and AI compute environments.
  • Assist with the operation of large-scale GPU clusters built on NVIDIA HGX, DGX, or similar accelerated computing platforms.
  • Troubleshoot common hardware, operating system, driver, firmware, networking, and storage issues.
  • Investigate degraded, unavailable, or unstable compute and GPU nodes using established troubleshooting procedures.
  • Escalate complex or unfamiliar issues to senior administrators or engineering teams when appropriate.
  • Assist with identifying recurring infrastructure issues and documenting technical findings.
  • Support bare-metal, virtualized, containerized, and cloud-hosted systems.
  • Maintain operational documentation, support procedures, runbooks, and troubleshooting guides.

GPU and Accelerated Computing Systems

  • Assist with the installation, configuration, validation, and troubleshooting of NVIDIA GPU drivers, firmware, and supporting system software.
  • Monitor and troubleshoot GPU health using tools such as nvidia-smi, DCGM, and NVIDIA Fabric Manager.
  • Identify and collect diagnostic information for GPU Xid errors, NVLink issues, PCIe errors, ECC events, GPU resets, and thermal or power-related issues.
  • Assist senior team members with troubleshooting multi-GPU communication and performance issues.
  • Perform post-maintenance validation of GPU systems following repairs or infrastructure changes.
  • Coordinate and track hardware replacement and RMA activities for GPUs, system boards, NICs, power supplies, and other server components.

Linux Systems Administration

  • Administer and support enterprise Linux systems.
  • Troubleshoot common boot, service, filesystem, memory, CPU, device-discovery, and operating-system issues.
  • Support Linux networking, storage mounts, authentication, permissions, SSH, DNS, and NTP.
  • Collect diagnostic information for system crashes, kernel issues, hardware errors, and out-of-memory conditions.
  • Follow established security and operating-system configuration standards.

Server Hardware, Firmware, and Out-of-Band Management

  • Support enterprise server platforms from vendors such as Dell, HPE, ASUS, Supermicro, NVIDIA, or similar manufacturers.
  • Troubleshoot common hardware issues involving memory, PCIe devices, GPUs, NICs, storage devices, power supplies, fans, system boards, and cabling.
  • Assist with BIOS, BMC, NIC, GPU, drive, and other component firmware updates.
  • Use out-of-band management tools to perform remote console access, power operations, hardware inventory collection, log collection, and system health checks.
  • Work with Data Center technicians during hardware replacements, cabling validation, break-fix activities, and post-repair testing.
  • Follow established procedures for firmware upgrades and hardware maintenance.

High-Performance Networking

  • Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments.
  • Support NVIDIA/Mellanox and other network adapters used in compute infrastructure.
  • Perform basic troubleshooting of link state, interface errors, MTU configuration, packet loss, VLANs, RDMA, RoCE, and InfiniBand connectivity.
  • Collect network diagnostic information and escalate complex fabric or switching issues to Network Engineering or senior team members.
  • Assist with validation following network changes, firmware upgrades, cable replacements, and new deployments.

Incident, Problem, and Change Management

  • Participate in incident response for production infrastructure events.
  • Assist with troubleshooting activities and provide clear technical updates during incidents.
  • Document troubleshooting steps, findings, actions taken, and resolution details.
  • Escalate incidents according to established support and escalation procedures.
  • Follow defined change-management processes, maintenance procedures, and rollback plans.
  • Work within established service-level agreements and operational processes.
  • Participate in a scheduled on-call rotation as required.

Collaboration and Development

  • Work closely with senior Systems Administrators and engineers to develop HPC infrastructure troubleshooting skills.
  • Follow established troubleshooting methods, documentation standards, and operational procedures.
  • Contribute to runbooks, knowledge-base articles, and support documentation.
  • Participate in technical reviews, troubleshooting sessions, and post-incident discussions.
  • Share technical findings and lessons learned with other team members.
  • Continue developing knowledge of Linux, GPU infrastructure, high-performance networking, storage, and automation.

 

QUALIFICATIONS REQUIRED (What Sets You Apart)

  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role. Equivalent hands-on experience, education, or technical training may also be considered.
  • Working knowledge of Linux operating systems and command-line administration.
  • Basic understanding of enterprise server hardware and common server components.
  • Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools.
  • Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting.
  • Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware.
  • Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms.
  • Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools.
  • Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies.
  • Ability to follow technical procedures and troubleshoot problems methodically.
  • Ability to recognize when an issue requires escalation and clearly communicate collected findings.
  • Strong written communication and documentation skills.
  • Ability to work collaboratively across technical teams.
  • Participate in an on-call rotation.

 

STRONGLY PREFERRED

  • Exposure to NVIDIA HGX, DGX, H100, H200, B100, B200, or similar GPU platforms.
  • Familiarity with nvidia-smi, DCGM, NVIDIA Fabric Manager, CUDA, or GPU diagnostic tools.
  • Exposure to InfiniBand, RDMA, RoCE, NVIDIA/Mellanox adapters, or NVIDIA UFM.
  • Familiarity with NVLink, NVSwitch, NCCL, or multi-GPU systems.
  • Experience working with enterprise server hardware from Dell, HPE, Lenovo, Supermicro, NVIDIA, or similar vendors.
  • Exposure to MAAS or other bare-metal provisioning technologies.
  • Familiarity with NFS, Lustre, BeeGFS, Spectrum Scale, Ceph, or other storage technologies.
  • Experience with monitoring tools such as Prometheus, Grafana, Datadog, DCGM Exporter, or similar platforms.
  • Experience with hardware diagnostics, vendor support cases, or RMA processes.
  • Familiarity with Git, Ansible, Terraform, or other infrastructure automation tools.
  • Familiarity with Jira Service Management or another ITSM platform.
  • Certifications such as RHCSA, CompTIA Linux+, Linux Foundation, NVIDIA, CCNA, or equivalent technical training.

 

WORKING CONDITIONS

  • Remote-first position with extensive computer-based work.
  • Work may involve coordinating remote troubleshooting and hands-on activities with Data Center technicians.
  • Fast-paced, customer-facing production environment supporting business-critical HPC and AI infrastructure.
  • Participation in a scheduled on-call and escalation rotation is required.
  • Occasional work during maintenance windows, weekends, or outside normal business hours may be required.
  • Additional responsibilities may be assigned based on operational requirements and company growth.

 

5C Data Centers is an equal opportunity employer.

At 5C Data Centers, we celebrate diversity and are committed to creating an inclusive environment where everyone can thrive.

5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.

 

 

#LI-AJ1

 

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

[Job-30394] Senior Fullstack Developer (React / Java / Node), Brazil

Full Time Brazil Software Development

General QA Engineer (gambling)

Full Time Slovakia Software Development

Senior Solution Engineer - Strategic Accounts Americas

Full Time United States Software Development

Java + AI Tools (Claude code) - USA

Full Time United States $133K - $140K per year Software Development

Jr Java Developer - USA

Full Time United States $112K - $113K per year Software Development

Outside Sales Engineer

Full Time United States Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 217,826+ Jobs in Systems Administrator

Answer easy questions

Answer easy questions

217,826+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified