Machine Learning Engineer, Speech - Joint Audio-Video Modeling
You will design and implement state-of-the-art speech and audio generation systems, focusing on joint audio-video modeling and production inference. You will also drive the end-to-end model development lifecycle, including data curation, experimental design, and performance optimization.