For Employers

Cudo Ventures

HPC Engineer

Posted 17 hours ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will be responsible for the end-to-end deployment, tuning, and automation of large-scale GPU infrastructure clusters. Additionally, you will provide operational support, documentation, and technical guidance to the service team and customers.


CUDO is delivering a large-scale GPU infrastructure project, taking it from hardware specification through to production clusters running real AI workloads. Getting there means more than installing GPUs. The infrastructure needs to be stable, tuned, monitored and built to operate reliably from day one.

As an AI infrastructure owner-operator, CUDO doesn't design something and hand it over. We own it and we run it. So the GPU clusters we build need to be specified properly, deployed cleanly, automated from day one and supported by processes our service team can operate.

Why this role exists

That project creates a very specific engineering challenge, and it is why we need an HPC Engineer now. Delivering GPU infrastructure at this scale takes hands-on engineering across every layer of the cluster: compute, interconnect, storage, scheduling and monitoring.

You will work alongside our Senior HPC Engineers to get these clusters built, tuned and running reliably. Just as importantly, you will help make sure the way we deliver them is repeatable, documented and ready for whatever CUDO builds next.

  • Shaping the build before the hardware arrives. Work with OEMs and vendors on bills of materials, data centre layouts and power and cooling specifications.
  • Deploying and tuning the clusters. Install and tune Linux, design and test InfiniBand and RoCE networking, configure SLURM partitions and priorities, and set up parallel file systems such as Ceph, Lustre, WEKA and VAST.
  • Automating the delivery. Use Ansible, Terraform and scripting to standardise deployment, configuration and maintenance, so every build is faster and more consistent than the last.
  • Making it operable. Build proactive monitoring with our SREs, so issues are caught and fixed before users feel them. Keep clusters secure, patched, resilient and highly available.
  • Handing over well. Write clear architecture, configuration and procedure documentation, train the service team, move routine work to them and act as their escalation point for AI and HPC incidents.
  • Advising customers. Support the sales team in pre-sales conversations, helping to architect solutions that CUDO can deliver and run.
  • Building for what comes next.  Track developments from NVIDIA, AMD and Intel and bring that knowledge into how we plan and build.

You know what it takes to make complex compute infrastructure work in production. You have been hands-on with AI or HPC environments at scale and are comfortable working across the cluster rather than within one narrow layer. You will bring:

  • Strong Linux administration, networking and storage skills
  • Experience deploying and tuning high-performance networking (InfiniBand, RDMA, RoCE)
  • Hands-on experience with HPC job schedulers, ideally SLURM
  • Experience with parallel and distributed file systems
  • In-depth knowledge of parallel, distributed and GPU computing
  • Scripting and automation skills (Bash, Python) and familiarity with Ansible, Terraform or similar
  • Experience with performance profiling, benchmarking and tuning of compute and I/O workloads
  • A solid grasp of security, patching, system hardening, high availability and ITSM practices
  • Clear written communication, especially procedures, architecture notes and incident reports
  • A capacity planning mindset: you think about what the infrastructure will need to handle next
  • A degree in Computer Science, Computer Engineering, Electrical Engineering or a related field, or equivalent hands-on experience

You may also have:
  • Managed very large AI and HPC clusters (hundreds to thousands of nodes), or run HPC in cloud or hybrid environments
  • Experience with vendor hardware validation, procurement planning and power and cooling design
  • Worked with containerisation in HPC (Apptainer, Singularity, Docker) and tools such as Spack, EasyBuild and Lmod
  • Knowledge of current AI and ML workflows, GPU-aware MPI and NVLink
  • Delivered AI or HPC clusters in cloud, research or academic settings
  • Relevant certifications such as RHCE, ITSM or vendor training
  • A Master's or PhD in a relevant STEM field
  • Reports to: Head of Engineering
  • Team: Infrastructure
  • Location: Remote
  • Travel: International travel as the work require

If you want to take AI infrastructure from bill of materials to production workloads, and own how it is built and run, we would like to hear from you. Tell us what you have built, the scale you have worked at and the problems you solved along the way.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Tupande AI Engineering Lead (Fixed-term)

Full Time Estonia, Kenya, Oman +5 more Software Development

Specjalist_ ds. sprzedaży AI (klient biznesowy)

Full Time Poland 10000 per month Software Development

Analista Programador/A Senior Full-Stack - Madrid (Remoto)

Full Time Spain €170 per day Software Development

Senior Data Engineer, Databricks

Full Time United States $135K - $170K per year Software Development

Staff Software Engineer (L6) / TLM - Developer Productivity - Platform Systems, AIMS Engineering

Full Time United States $600K - $1066K per year Software Development

Senior Data Engineer - Data Platform

Full Time India Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 217,817+ Jobs in Software Development

Answer easy questions

Answer easy questions

217,817+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified