You will operate and maintain the language model benchmarking pipeline while serving as the primary technical contact for AI lab customers. This involves debugging complex issues, communicating results, and ensuring the accuracy and integrity of all published benchmark data.
Lead the initiation of robotics coverage by developing benchmarking methodologies, leaderboards, and strategic analysis for AI-driven robotics. Collaborate with leading robotics labs and AI companies to evaluate VLAs, world models, and full robotic systems.
Manage the end-to-end media generation benchmarking pipeline, including running image and video evaluations and managing human preference studies. Serve as the technical point of contact for model providers to communicate results and explain methodologies.