CUDO is delivering a large-scale GPU infrastructure project, taking it from hardware specification through to production clusters running real AI workloads. Getting there means more than installing GPUs. The infrastructure needs to be stable, tuned, monitored and built to operate reliably from day one.
As an AI infrastructure owner-operator, CUDO doesn't design something and hand it over. We own it and we run it. So the GPU clusters we build need to be specified properly, deployed cleanly, automated from day one and supported by processes our service team can operate.
Why this role exists
That project creates a very specific engineering challenge, and it is why we need an HPC Engineer now. Delivering GPU infrastructure at this scale takes hands-on engineering across every layer of the cluster: compute, interconnect, storage, scheduling and monitoring.
You will work alongside our Senior HPC Engineers to get these clusters built, tuned and running reliably. Just as importantly, you will help make sure the way we deliver them is repeatable, documented and ready for whatever CUDO builds next.
- Shaping the build before the hardware arrives. Work with OEMs and vendors on bills of materials, data centre layouts and power and cooling specifications.
- Deploying and tuning the clusters. Install and tune Linux, design and test InfiniBand and RoCE networking, configure SLURM partitions and priorities, and set up parallel file systems such as Ceph, Lustre, WEKA and VAST.
- Automating the delivery. Use Ansible, Terraform and scripting to standardise deployment, configuration and maintenance, so every build is faster and more consistent than the last.
- Making it operable. Build proactive monitoring with our SREs, so issues are caught and fixed before users feel them. Keep clusters secure, patched, resilient and highly available.
- Handing over well. Write clear architecture, configuration and procedure documentation, train the service team, move routine work to them and act as their escalation point for AI and HPC incidents.
- Advising customers. Support the sales team in pre-sales conversations, helping to architect solutions that CUDO can deliver and run.
- Building for what comes next. Track developments from NVIDIA, AMD and Intel and bring that knowledge into how we plan and build.
You know what it takes to make complex compute infrastructure work in production. You have been hands-on with AI or HPC environments at scale and are comfortable working across the cluster rather than within one narrow layer. You will bring:
- Strong Linux administration, networking and storage skills
- Experience deploying and tuning high-performance networking (InfiniBand, RDMA, RoCE)
- Hands-on experience with HPC job schedulers, ideally SLURM
- Experience with parallel and distributed file systems
- In-depth knowledge of parallel, distributed and GPU computing
- Scripting and automation skills (Bash, Python) and familiarity with Ansible, Terraform or similar
- Experience with performance profiling, benchmarking and tuning of compute and I/O workloads
- A solid grasp of security, patching, system hardening, high availability and ITSM practices
- Clear written communication, especially procedures, architecture notes and incident reports
- A capacity planning mindset: you think about what the infrastructure will need to handle next
- A degree in Computer Science, Computer Engineering, Electrical Engineering or a related field, or equivalent hands-on experience
You may also have:
- Managed very large AI and HPC clusters (hundreds to thousands of nodes), or run HPC in cloud or hybrid environments
- Experience with vendor hardware validation, procurement planning and power and cooling design
- Worked with containerisation in HPC (Apptainer, Singularity, Docker) and tools such as Spack, EasyBuild and Lmod
- Knowledge of current AI and ML workflows, GPU-aware MPI and NVLink
- Delivered AI or HPC clusters in cloud, research or academic settings
- Relevant certifications such as RHCE, ITSM or vendor training
- A Master's or PhD in a relevant STEM field
- Reports to: Head of Engineering
- Team: Infrastructure
- Location: Remote
- Travel: International travel as the work require
If you want to take AI infrastructure from bill of materials to production workloads, and own how it is built and run, we would like to hear from you. Tell us what you have built, the scale you have worked at and the problems you solved along the way.