You will own the IP/Ethernet infrastructure connecting the global GPU cloud, including DCI, WAN, and internet edge networks. Additionally, you will build automation to transition network operations from manual tickets to intent-driven policy as code.
The role involves designing and maintaining CI/CD and MLOps pipelines while managing cloud-native infrastructure using Kubernetes and Docker. You will also act as a technical lead for incident management and ensure high availability and security across AI-powered platforms.
You will architect and maintain the high-performance hardware foundation for AI-native cloud computing by integrating GPU device plugins and optimizing network stacks. Additionally, you will manage automated hardware remediation pipelines and tune kernel-level parameters to ensure low-latency, high-bandwidth performance for distributed AI workloads.
Architect and scale a high-cardinality telemetry infrastructure to monitor massive-scale AI hardware and software performance. Develop automated diagnostic tools and alerting pipelines to ensure real-time detection and resolution of system bottlenecks.
The Staff Backend Engineer will co-own the end-to-end design and implementation of a globally distributed, multi-tenant Model-as-a-Service platform. This role involves driving technical strategy, optimizing performance for high-throughput AI inference, and ensuring robust reliability and billing accuracy.
Design and implement advanced batch scheduling architectures to optimize AI workload placement and resource utilization. Collaborate with hardware teams to integrate scheduling layers with bare-metal infrastructure and manage complex accelerator requests.
You will contribute to the development of the NeoCloud SRE platform by building monitoring, storage, and automation components under the guidance of senior engineers. This involves writing production code, implementing observability instrumentation, and participating in on-call rotations to maintain system reliability.
You will architect and build a highly available, automated control plane for the NeoCloud platform to manage thousands of Kubernetes clusters. This involves developing custom operators, optimizing etcd performance, and ensuring seamless multi-cluster federation for AI workloads.
You will architect and maintain high-performance storage solutions to support AI model training and inference workloads. This involves managing container storage interfaces, optimizing I/O throughput, and ensuring seamless data delivery for GPU-intensive applications.
You will deploy and operate high-performance parallel storage systems to support AI training and inference workloads. Additionally, you will instrument storage telemetry and partner with the platform team to develop predictive models for storage-fault detection.
You will design, deploy, and operate the control plane for AI-operated GPU cloud infrastructure, ensuring high availability and automated remediation. You will manage production Kubernetes clusters at scale, implementing topology-aware scheduling and automated failure handling for GPU workloads.
You will lead the end-to-end architecture design of AI Data Center Networks, high-performance Data Center Interconnects, and global backbone networks. Additionally, you will collaborate with engineering teams to implement and optimize hyperscale cluster network standards and configurations.
The Senior Power Consultant will manage wholesale and retail energy market activities, including power scheduling, demand response strategy, and hedging. They will also lead power contract negotiations and provide market intelligence to optimize operational costs across global mining facilities.