Please mention DailyRemote when applying
About Civo:
Civo is a high-performance neocloud provider purpose-built for the demands of modern AI, high-performance computing (HPC), and cloud-native infrastructure. We eliminate legacy cloud overhead to deliver ultra-low-latency compute, bare-metal GPU performance, and streamlined Kubernetes orchestration at scale.
Purpose-designed for AI engineering teams, enterprises, and research institutions, Civo delivers direct access to cutting-edge NVIDIA GPU clusters, high-speed fabrics, and parallel storage systems required to train, fine-tune, and deploy foundation models efficiently. We combine high-density infrastructure with predictable pricing and maximum compute throughput, empowering organizations to scale AI workloads without the complexity or cost bloat of traditional hyperscalers.
About the Role:
As a Senior Solution Engineer – GPU & AI Infrastructure, you will serve as the primary technical architect for Civo’s large-scale AI and high-performance computing (HPC) customer initiatives. You will be responsible for designing state-of-the-art NVIDIA GPU clusters tailored for training and inferencing massive foundation models.
In this role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare-metal and Kubernetes-based orchestrations across cutting-edge NVIDIA Blackwell architectures (e.g., B300 and GB300NVL) using ultra-low-latency InfiniBand and high-speed RoCE networking fabrics.
Responsibilities:
Solution Design & Architecture
System Design Documents: Author comprehensive High-Level Design (HLD) and Low-Level Design (LLD) documentation for enterprise-scale GPU supercomputing clusters.
Bill of Materials (BOM): Generate detailed BOMs covering compute nodes, NVLink switches, network fabrics, transceivers/cabling, liquid/air cooling requirements, power distribution, and high-performance storage.
GPU Cluster Topology: Architect scale-up (NVLink/NVSwitch) and scale-out network topologies (Fat-Tree, Rail-Optimized) for NVIDIA Blackwell platforms, specifically B300 and GB300NVL rack-scale architectures.
Fabric & Networking Engineering: Design high-throughput, low-latency networking architectures utilizing both InfiniBand (e.g., NDR/X800) and RoCE / RoCEv2 (e.g., NVIDIA Spectrum-X / Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing).
Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native / Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow).
Storage Integration: Architect high-bandwidth parallel storage solutions utilizing GPUDirect Storage (GDS) and enterprise AI file systems (e.g., VAST Data).
Technical Sales Support & Customer Engagement
Partner with Civo’s sales and commercial teams as the technical lead for high-value AI infrastructure opportunities.
Engage directly with customer CTOs, Chief AI Officers, infrastructure leads, and ML engineers to evaluate technical requirements, compute sizing, and fabric choices.
Lead deep-dive architectural workshops and technical presentations on Civo's bare-metal GPU and managed Kubernetes offerings.
Produce precise technical proposals and lead responses to complex RFPs/RFIs regarding AI infrastructure.
Proof-of-Concept (PoC) & Benchmarking
Architect and oversee Proof-of-Concept (PoC) deployments to validate real-world performance for customer workloads.
Benchmark cluster performance using industry-standard tools (NCCL tests, GPUDirect RDMA latency/bandwidth, MLPerf, Megatron-LM benchmarks).
Address network congestion, fabric routing, and thermal/power optimization during validation phases.
Product & Ecosystem Collaboration
Serve as the bridge between enterprise AI clients, hardware vendors (NVIDIA, network OEMs), and Civo’s internal platform engineering team.
Provide continuous feedback to product teams on market trends, hardware platform demands, and feature requirements for AI/GPU orchestration.
Key Results/Objectives:
Technical Wins: Achieve high technical win rates on large-scale AI/GPU cluster sales opportunities.
Design Excellence: Successfully deliver complete, peer-reviewed HLDs, LLDs, and BOMs within target deal timelines.
Customer Satisfaction: Achieve successful PoC completion and sign-off for enterprise clients scaling AI workloads on Civo infrastructure.
Requirements:
5+ years in a Solution Architecture, Systems Engineering, or Technical Pre-Sales role focused on high-performance cloud, HPC, or AI infrastructure.
Bachelor’s degree in Computer Science, Electrical Engineering, Systems Engineering, or equivalent practical experience.
NVIDIA GPU Architecture: Deep hands-on knowledge of NVIDIA HGX/DGX platforms, NVLink/NVSwitch fabrics, and Blackwell architectures (B300, GB300NVL, GB200 NVL72/NVL36).
High-Speed Networking: Expert-level knowledge of cluster fabric topologies:
InfiniBand: Quantum-2 / Quantum-X800, Subnet Management, Adaptive Routing.
RoCE / RoCEv2: Spectrum-X / Spectrum-4 Ethernet switches, PFC, ECN, RoCE configuration, and optimization.
GPU Direct Technologies: GPUDirect RDMA (GDR) and GPUDirect Storage (GDS).
Orchestration & Platforms: Proficiency in deploying and optimizing GPU workloads on:
Kubernetes: Container networking (CNI), NVIDIA GPU Operator, RDMA Shared Device Plugin, MPI Operator.
Bare-Metal: Slurm, Ansible, Terraform, PyTorch/NCCL environment tuning.
Documentation Skills: Demonstrated experience creating enterprise-grade HLDs, LLDs, network rack diagrams, and itemized BOMs.
Power & Thermal Awareness: Familiarity with high-density datacenter environments, liquid cooling technologies (Direct-to-Chip, CDU/liquid loop setups), and power delivery constraints for 100kW+ per rack deployments.
Strong technical leadership and presentation skills, with the ability to articulate complex network and hardware tradeoffs to executive stakeholders.
Problem-solving mindset capable of diagnosing complex hardware-software interaction bottlenecks in distributed training/inference setups.
Must be UK based.
Nice to Have:
NVIDIA Certified Professional: AI Infrastructure (NCP-AII).
NVIDIA Certified Professional: AI Networking (NCP-AIN).
NVIDIA Certified Professional: InfiniBand (NCP-IB).
NVIDIA Certified Associate / Professional: AI Workload Deployment & Cloud Native.
Why Join Civo?
Competitive compensation and benefits package.
4-day week company (unless attending an event).
Uncapped holiday.
Remote work environment with flexibility and autonomy.
Collaborative and inclusive culture that values diversity and creativity.
Opportunity to work with a dynamic and innovative team in the fast-growing cloud industry.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!