You will evolve the observability stack for logs, metrics, and traces while ensuring reliable telemetry across cloud and on-premise dataplanes. Additionally, you will lead incident response, define SLOs, and optimize telemetry costs to maintain high system availability.
You will own and evolve the inference platform, managing model execution for both real-time APIs and large-scale batch processing. This involves optimizing performance, cost, and reliability across cloud and customer-hosted Kubernetes environments.
You will build and maintain the ML systems platform, focusing on distributed training, model governance, and reproducible release pipelines. This involves developing compute primitives, managing data formats, and ensuring all production models are fully auditable and lineage-tracked.
You will define and execute an applied research agenda while building and optimizing end-to-end machine learning pipelines. Additionally, you will collaborate with cross-functional teams to integrate frontier research into scalable, real-world AI applications.