You will architect and build major components of the AI cloud platform, including GPU and network virtualization stacks and global management planes. You will also lead technical direction across teams, mentor engineers, and ensure the reliability and scalability of the infrastructure.
You will act as the first line of support for customers building AI solutions, managing inbound tickets and resolving technical issues. You will also collaborate with engineering and product teams to improve platform offerings and ensure high customer satisfaction.
Design and build production AI agent systems to diagnose, investigate, and remediate infrastructure issues across a massive GPU fleet. Develop the underlying distributed services, orchestration frameworks, and knowledge graphs that power these autonomous agents.
You will act as a technical partner to strategic customers, optimizing inference engines and guiding fine-tuning pipelines for production AI models. Additionally, you will influence the product roadmap by surfacing field insights and ensuring successful platform adoption for complex use cases.
You will design and build production AI agent systems to diagnose and remediate infrastructure issues across a massive GPU fleet. Additionally, you will develop the orchestration frameworks, knowledge graphs, and retrieval systems that power these autonomous agents.
Design and build production AI agent systems to diagnose, investigate, and remediate infrastructure issues across a large-scale GPU fleet. Develop the distributed services, orchestration frameworks, and knowledge graphs that power these autonomous infrastructure agents.
The manager will oversee international cloud compute supplier relationships, manage contract lifecycles, and run RFP processes for new capacity. They will also develop market intelligence and negotiation strategies to optimize costs and ensure supply chain efficiency for the global GPU fleet.
You will act as a technical partner to strategic customers, optimizing inference engines and guiding fine-tuning pipelines for production AI models. Additionally, you will influence the product roadmap by surfacing field insights and ensuring successful platform adoption.
United States$200K - $250K per year5-10 yrs expProduct
Own and improve the end-to-end qualification process for new compute capacity to ensure it meets technical standards. Coordinate with engineering teams to validate hardware, networking, and operational resilience before recommending go/no-go decisions.
Maintain user-facing services and production systems using automation and sound engineering principles. Build and scale AI infrastructure using Ansible, Terraform, and Kubernetes while managing on-call rotations and monitoring systems.
Design and implement a scalable observability platform for metrics, logs, and traces to monitor system performance and GPU utilization. Develop automated monitoring, alerting systems, and custom infrastructure-as-code tools to enhance distributed tracing and incident response.
Provide analytical support for infrastructure strategy, focusing on capacity planning, compute sourcing, and site selection. Build dashboards and frameworks to monitor deployments and streamline operational workflows across Engineering and Finance.
Design and build infrastructure for model customization and evaluation, including backend services and job orchestration platforms. Collaborate cross-functionally to improve platform reliability, observability, and deployment tooling.
Design and deliver multi-petabyte storage systems and high-performance parallel filesystems optimized for AI training and inference workloads. Build Kubernetes-native storage operators and optimize end-to-end data paths to achieve high throughput and low latency.
Partner with AI research and engineering leadership to define and execute hiring strategies for world-class research teams. Lead full-cycle recruiting for specialized AI talent across academia and industry while refining technical interview processes.
You will optimize and fine-tune GPU code for better performance and scalability while collaborating with cross-functional teams to integrate GPU-accelerated solutions. Staying up-to-date with advancements in GPU programming techniques is also a key responsibility.