AI VFX
Senior Applied Scientist - Multimodal
External listing
External listing from the employer
Originally published by Flawless through its Ashby careers source. Applications go to the employer's website, and AIMovieJobs is not the hiring employer.
View original listingQuick facts
LondonLocationHybridWorkplaceFull-timeJob typeCompetitive salaryCompensationAI VFXCategoryScienceDepartmentJul 20, 2026PostedJul 26, 2026Last verifiedJan 22, 2027Expires
Role overview
"The AI company that's revolutionizing Hollywood"Flawless is transforming Hollywood with assistive AI. Our tools empower filmmakers to edit, localize, and refine performances while preserving artistic intent.
What you'll do
- Develop repeatable, scalable audio/video dataset curation pipelines and lip sync model training workflows across multiple datasets
- Train, fine-tune, and manage audio/video and lip sync model variants as model dependencies, data, and architectures evolve
- Incorporate new datasets and model updates as they become available
- Design, automate, and maintain audio/video datasets and lip sync metric testing pipelines
- Generate new quantitative and qualitative metrics to evaluate audio/video and lip sync quality
- Produce comparisons, visualizations, and analyses to inform research and product decisions
- Partner closely with audio/video and lip sync researchers to support ongoing and future research initiatives
- Validate audio/video and lip sync quality to improve out-of-the-box approval rates and reduce downstream cost and iteration time
- Collaborate with Science, Engineering, and Product teams to align research outputs with company goals
What the employer is looking for
- MSc OR PhD + Industry experience working in the domains of Audio processing, 3D Computer Vision, Speech Synthesis, Computer Graphics, or other multimodal related fields such as text/audio, or audio/visual.
- Proficiency in Python, with a strong foundation in computer science and problem-solving.
- Expertise with deep learning frameworks (PyTorch) and vision tools (OpenCV).
- A strong product mindset — motivated by building systems that deliver tangible value to users, not just technical novelty.
- Comfortable working at both the algorithmic and implementation levels, from model design and optimisation to large-scale data processing and integration in production systems.
- High degree of proficiency in math and statistical methods for signal processingExperience with audio-visual learning, multimodal fusion, and/or audio-driven face animation
- Experience with speech processing and detection, such as dialog/speaker detection, speaker separation, and speech synthesis with deep neural networks
- Outstanding communication skills for collaboration with scientists, research/ML engineers, and VFX artists
- Absolutely! Research shows that women and underrepresented groups often hesitate to apply unless they meet every qualification, but at Flawless, we actively work to break down those barriers. We believe diverse perspectives, experiences, and backgrounds make us stronger, and we are committed to supporting and elevating underrepresented talent. If you're excited about the role, share our values, and believe you can contribute meaningfully, we encourage you to apply—even if you don’t meet every single requirement. Your unique skills and perspective matter, and we’d love to hear from you ❤️
Preferred qualifications
- Demonstrable research experience with a strong publication record in major 3D Computer Vision, Speech Processing, and Computer Graphics venues and journals (e.g., CVPR, SIGGRAPH, NeurIPS)
- Experience developing multi-modal systems that integrate audio, text, and visual inputs.
- Experience working with cross-functional teams
- Experience with generative and cross-domain attention models for audio/visual-based speech applications