The Principal Hardware Systems Engineer will design, deploy, and manage large-scale AI cloud hardware infrastructure. This role involves direct engagement with customers to define requirements and providing technical leadership to ensure high performance and reliability of AI workloads.
Member of Technical Staff, Hardware Engineering
Location: Bay Area/ Seattle/ Remote
Reports to: CTO
About Neo Cloud
Neo Cloud is a fast growing next-generation AI neocloud, founded by ex NVIDIA, CoreWeave and Intel engineering leaders. We offer you the opportunity to build and operate next generation AI infrastructure that powers the world's most demanding AI workloads for the largest AI labs. We are rethinking how to run large scale AI infra for instance by leveraging agentic AI to build digital twins to derisk our physical deployments. Our agents are deploying and monitoring our fleet to maximize the uptime of our infra.
About the Role
We're looking for a Principal Hardware Systems Engineer to design, deploy and manage Neo Cloud's AI hardware. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations — translating what AI/ML customers will need in the future, into a hardware design and operation at massive scale.
Key Responsibilities
Identifying customer requirements
Engage directly with customers, solutions architects, and product teams to understand future hardware requirements for AI/ML workloads.
Extrapolate from current usage patterns and industry trends to anticipate future requirements.
Partner with vendors and drive their roadmaps.
Hardware design, deployment, and operations:
Design, implement, and operate AI cloud hardware.
Architect, evaluate, and Deploy NVIDIA GPU infrastructure, including Blackwell-based B200, GB200, GB300, DGX and NVL rack-scale systems.
Design rack-scale infrastructure encompassing NVLink and NVSwitch fabrics, Infiniband or Spectrum-X Ethernet with RoCE for scale-out networking; power delivery, liquid cooling, cabling and serviceability.
Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions.
Architect for the specific demands of AI workloads focusing on performance, security, resiliency and cost.
Design and implement consistent and efficient hardware operation.
Take end-to-end ownership of hardware from design, supply chain planning, deployment, operation and ultimately deprecation.
Ensure security and compliance.
Strong debugging skills and analytical skills to identify failures across a variety of electrical, thermal, optical, mechanical and software failures.
Technical leadership:
Strong bias to action and resolution of technical decisions.
Set technical direction and best practices for the hardware organization; author and review design documents for significant architectural changes.
Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
Influence technical strategy across adjacent teams.
Qualifications
10+ years of professional hardware engineering experience building and operating hardware at massive scale.
Direct experience with AI-focused hardware deployments across different hardware vendors, including NVIDIA GPU platforms.
Experience with NVIDIA BLackwell, GB200 or GB300 NVL72.
Experience with NVLink, NVSwitch, InfiniBand or Spectrum-X Ethernet, ConnectX networking, and BlueField DPUs.
Deep system level expertise.
Proven experience operating large-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
Strong systems programming skills.
Excellent written and verbal communication skills.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”