The Compute Procurement Lead is responsible for sourcing global GPU capacity and managing relationships with hyperscalers, data centers, and chip vendors. This role involves technical evaluation of infrastructure, commercial negotiation of multi-billion dollar spend, and conducting on-site assessments.
You will embed with customer teams to onboard, tune, and debug large-scale AI training and inference workloads on GPU clusters. You are also responsible for building automation and infrastructure to ensure the reliability and performance of these distributed systems.
You will own the backend systems responsible for capacity quoting, reservation, funding, and billing across the marketplace. Additionally, you will manage the full API surface and infrastructure operations, including metering and customer-facing dashboards.
You will build and operate the control plane for a large-scale GPU fleet, managing the automated lifecycle of clusters from bare metal to customer-ready. This includes maintaining Kubernetes and Postgres systems while scaling infrastructure to support thousands of nodes.
You will design and build high-performance systems across storage, networking, virtualization, and container runtimes. Additionally, you will participate in on-call rotations to respond to production incidents.
Own provider relationships and execution, managing onboarding, capacity rollouts, and hardware quality issues. Act as the incident commander for major provider-side failures, coordinating responses between SREs, providers, and customers.
You will lead financial planning, performance management, and deal execution for the company's global GPU compute portfolio. This includes building financial architecture, managing lender relationships, and overseeing capital allocation for infrastructure projects.
You will build and own the end-to-end partnership lifecycle, including sourcing, negotiating, and managing relationships with GPU providers, AI labs, and VC funds. You will also drive cross-functional coordination to ensure partnership agreements are successfully integrated into the company's infrastructure and sales operations.
You will design, operate, and debug large-scale GPU infrastructure for distributed AI training and inference. You will also serve as the primary technical partner for customers, ensuring the reliability and performance of high-speed interconnects and compute clusters.