You will architect and implement large-scale speech models while driving the model, data, and evaluation flywheel for voice conversion and controllable TTS tasks. Additionally, you will define data strategies, design rigorous evaluation metrics, and ensure production-ready performance for safe and steerable AI systems.
Cantina
9 Remote Job Openings at Cantina
Design and build core services and platform components to accelerate product development across the engineering organization. Oversee DevOps initiatives, manage infrastructure, and improve developer experience through optimized backend systems.
Machine Learning Engineer, Speech - Joint Audio-Video Modeling
Cantina
·
Full Time
·
6 days ago
Cantina
You will design and implement state-of-the-art speech and audio generation systems, focusing on joint audio-video modeling and production inference. You will also drive the end-to-end model development lifecycle, including data curation, experimental design, and performance optimization.
Design and scale inference infrastructure for generative audio models including TTS and ASR to ensure low-latency and reliability. Bridge the gap between research and production by optimizing model serving paths and automating CI/CD pipelines.
Lead the strategic direction and production of high-impact performance ads across platforms like Meta, TikTok, and YouTube to drive user acquisition. Manage internal creative teams and external agencies while utilizing AI tools to scale iterative, data-driven ad concepts.
Architect and build a shared-code strategy using Kotlin Multiplatform to power experiences across Android, iOS, and web. Design shared modules for networking, data persistence, and business logic while shipping polished Compose Multiplatform UIs.
Lead the development of high-performance Android features, including AI-driven capabilities and immersive custom UIs. Collaborate with product and design teams to architect scalable MVVM structures and optimize media pipelines for a social AI platform.
You will build and maintain scalable data pipelines for large video generation models, including ingestion, preprocessing, and dataset curation. Additionally, you will design annotation workflows and collaborate with research teams to optimize model-driven filtering systems.
The engineer will be responsible for designing, implementing, fine-tuning, improving, and debugging image AI models that power lifelike AI bots, focusing on generative image and machine vision services. This includes evaluating new research, developing and deploying new pipelines, monitoring production issues, and optimizing models for quality and latency.