Site Reliability Engineer
Design and operate reliable infrastructure for large-scale AI training and inference workloads. Automate operational workflows and build monitoring and incident-response practices to ensure system scalability and observability.