AI Sound & Music
Staff Research Engineer - Multimodal Generative Modelling
Originally published by Synthesia through its Ashby careers source. Applications go to the employer's website, and AIMovieJobs is not the hiring employer.
View original listingQuick facts
Role overview
As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Today over 60,000 businesses rely on our platform, and the next leap in what we can offer them depends on models that combine text, audio, and video into a single real-time interactive experience. You'll work directly with our voice lead and collaborate tightly with our video teams and other senior members of the org. Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
Preferred qualifications
- Experience with real-time or streaming architectures.
- Familiarity with state-of-the-art architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders.
- Excellence in one or more of the following modalities: voice, text, video.
- Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).