You will build and improve the inference layer of the Gcore platform by integrating frameworks like vLLM and TensorRT-LLM. Additionally, you will optimize model performance, including latency, throughput, and GPU utilization, while debugging complex issues across the stack.
You will build and maintain Python services and APIs for managing GPU virtual machines and bare-metal servers within a cloud infrastructure. Additionally, you will automate provisioning workflows, integrate hardware into Kubernetes, and investigate production issues across the control plane.
Develop and maintain RESTful APIs for managing GPU clusters and cloud infrastructure while bridging the gap between development and operations. You will participate in the full lifecycle management of compute products and design architectures focused on high availability and scalability.
Design and maintain the Traffic Management System API and agents to manage global routing and traffic-steering. Develop BGP and Anycast routing capabilities while ensuring system reliability through automated failover and observability.
Act as the final escalation point for complex Cloud infrastructure issues, diagnosing and resolving advanced incidents across compute, storage, and networking. Coordinate high-severity incident resolution with Engineering and DevOps teams while mentoring L1 and L2 engineers.
Lead the design, deployment, and maintenance of edge network infrastructure to ensure high availability and security. Mentor engineers and implement automation for system provisioning, monitoring, and configuration management.