Senior Site Reliability Engineer
You will scale and maintain a global GPU fleet while building robust Kubernetes infrastructure to ensure high availability. Additionally, you will automate operational tasks and develop observability tools to detect, survive, and recover from infrastructure failures.