Design and implement zero-trust security architecture for a GPU cloud platform, ensuring robust protection across identity, networking, and infrastructure layers. You will also develop automation workflows for threat detection, incident response, and vulnerability management to maintain a secure-by-design environment.
Install, validate, and maintain regional and core GPU cloud deployments across various datacenter environments. This role bridges the gap between physical infrastructure, networking, and platform operations to ensure production readiness.
Design, build, and operate the AI storage layer for large-scale GPU infrastructure across edge and core deployments. Optimize storage throughput and latency for distributed inference, fine-tuning, and training workloads.
Design, implement, and operate high-performance AI fabric and datacenter network infrastructure to support distributed GPU workloads. Act as the primary architectural authority for networking, driving automation, reliability, and cross-functional collaboration across the organization.
Design and build scalable observability platforms for large-scale GPU cloud infrastructure and edge deployments. Lead the delivery of major telemetry initiatives while collaborating with cross-functional teams to ensure platform reliability and performance visibility.
Provide advanced technical support for customers running workloads on GPU cloud platforms and serve as a technical escalation point for complex incidents. Collaborate with engineering teams to improve platform reliability, perform root cause analysis, and develop automation tools to streamline support workflows.