Partner with ML engineering teams to diagnose and resolve complex distributed systems and infrastructure issues in production. Drive long-term reliability improvements by building internal tooling, automation, and documentation.
Lightning AI
7 Remote Job Openings at Lightning AI
Design, build, and operate large-scale GPU infrastructure platforms to minimize incidents and enable AI features. Collaborate across engineering teams to handle break/fix operations, incident response, and automation to reduce manual toil.
Lead complex, cross-functional infrastructure programs spanning compute, network, storage, and datacenter operations. Drive the execution of high-impact initiatives, including hardware rollouts and capacity expansions, while managing risks and dependencies.
Design and deploy scalable spine-leaf network architectures and high-performance Ethernet fabrics to support GPU clusters and AI workloads. Develop automation and Infrastructure-as-Code solutions while optimizing traffic flows for AI training and inference.
Manage full-cycle recruiting for AI software engineering, infrastructure, and product teams. Focus on building a diverse candidate pipeline and improving scalable, equitable interview processes.
Design and operate large-scale GPU infrastructure platforms to minimize incidents and enable customer features. Collaborate across engineering teams to automate operational workflows and participate in an on-call rotation.
This is a general talent community expression of interest rather than a specific role. Candidates are invited to submit their information to be considered for future opportunities across the company's global hubs.