Machine Learning Engineer - Inference
Develop user-facing APIs, SDKs, and tools to streamline access to AI infrastructure. Collaborate with research and product teams to translate complex ML workflows into scalable and secure abstractions.
Develop user-facing APIs, SDKs, and tools to streamline access to AI infrastructure. Collaborate with research and product teams to translate complex ML workflows into scalable and secure abstractions.
Design and implement custom GPU/accelerator kernels to maximize performance for next-generation AI workloads. Collaborate with researchers to translate algorithmic advances into efficient, production-ready code.
Develop pipelines for post-training tasks including fine-tuning, evaluation, and model compression. Implement scalable systems for model deployment and optimization while collaborating with researchers to validate results in production.
The primary mission involves designing and optimizing large-scale pre-training systems to power generative AI models. This includes building scalable pre-training pipelines, implementing distributed training strategies across hardware, and developing monitoring systems for reliability.