Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
Design, implement, and operate high-performance AI fabric and datacenter network infrastructure to support distributed GPU workloads. Act as the primary architectural authority for networking, driving automation, reliability, and cross-functional collaboration across the organization.
Location & work modality: Remote
Type of Contract: Full time or Contract
About Radian Arc
We’re specialists in outcome-optimized AI infrastructure - deploying, orchestrating and monetizing GPU compute where data, users and demand actually meet: inside telco networks, at the edge, and in core data centers.Not a generic AI platform. Not a consultancy. We’re the bridge between raw silicon and real-world results.
What impact you will have
Design, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport.
As the first dedicated networking role in the organization, the Staff Network Engineer combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across design, deployment, troubleshooting, automation, and operational improvement.
The Staff Network Engineer owns the long-term technical direction and operational strategy for Radian Arc’s AI interconnect networks, designing scalable GPU fabrics and ensuring predictable low-latency performance across distributed training and inference workloads. The role includes designing large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuring operational excellence across global deployments, from hyperscale datacenters to smaller edge locations.
You will collaborate closely with platform, compute, storage, observability, and operations teams to ensure networking is deeply integrated into the overall infrastructure architecture. This role also acts as the senior escalation point for complex networking incidents, driving deep technical investigations and systemic improvements that increase reliability, latency consistency, and operational maturity across the platform.
Because this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards, cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers.
We expect to hire more than one person for this role so providing you have experience in designing AI Infrastructure networking solutions at scale please do not hesitate to apply if your experience does not match the full scope of the position.
What you’ll do
AI Fabric & HPC Networking
○ Leaf-spine
○ Fat-Tree
○ Rail architectures
○ Multi-plane
○ RDMA
○ RoCE
○ High-bandwidth east-west fabrics
○ Spectrum-X
Datacenter Networking
○ Network bridges
○ Routing stacks
○ Overlay networking systems
Technologies include:
Security & Edge Connectivity
Technologies include:
Inter-Datacenter Networking
Technologies include:
Engineering Execution & Delivery
Operational Excellence & Reliability
○ Provisioning
○ Configuration management
○ Monitoring
○ Lifecycle management
Cross-Functional Collaboration
Technical Stack
Datacenter Networking
Routing & Control Plane
Security
Transport & Backbone
AI Networking
What you'll need
Core Experience
○ BGP
○ OSPF
○ ECMP
○ EVPN / VXLAN
Advanced AI Fabric Networking Expertise
The candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads.
This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems.
Relevant expertise includes:
○ NCCL stalls
○ RDMA congestion
○ Fabric hotspots
○ Packet loss impacting distributed training
The candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, andhigh-throughput AI workloads.
Systems & Troubleshooting
○ Hardware
○ Firmware
○ Kernel networking
○ Distributed application communication layers
Automation
Leadership
What we offer
Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.
Our inclusive responsibility
Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,770+ Jobs in Network Engineer
Answer easy questions
216,770+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”