Partner with ML engineering teams to diagnose and resolve complex distributed systems and infrastructure issues in production. Drive long-term reliability improvements by building internal tooling, automation, and documentation.
Design, build, and operate large-scale GPU infrastructure platforms to minimize incidents and enable AI features. Collaborate across engineering teams to handle break/fix operations, incident response, and automation to reduce manual toil.
Lead complex, cross-functional infrastructure programs spanning compute, network, storage, and datacenter operations. Drive the execution of high-impact initiatives, including hardware rollouts and capacity expansions, while managing risks and dependencies.
Design and operate large-scale GPU infrastructure platforms to minimize incidents and enable customer features. Collaborate across engineering teams to automate operational workflows and participate in an on-call rotation.
This is a general talent community expression of interest rather than a specific role. Candidates are invited to submit their information to be considered for future opportunities across the company's global hubs.