AI Video Generation
Research Engineer
Originally published by Hedra through its Ashby careers source. Applications go to the employer's website, and AIMovieJobs is not the hiring employer.
View original listingQuick facts
Role overview
As a Research Engineer on our Physical AI team, you will lead pre-training and post-training on action-conditioned world models, working hand-in-hand with industrial partners to close the loop between generative AI and physical systems. Responsibilities:Design, implement, and run pre-training and post-training pipelines for action-conditioned world models and vision-language-action (VLA) modelsDevelop and refine training methodologies, including fine-tuning, reinforcement learning, and large-scale multimodal learningDesign and generate training and evaluation datasets from simulation, including environment setup, domain randomization, and sim-to-real transfer strategiesBuild distributed training infrastructure using PyTorch, FSDP, and DeepSpeedWork with multimodal data pipelines involving video, sensory inputs, and action sequencesEvaluate model performance using both benchmark datasets
What you'll do
- Design, implement, and run pre-training and post-training pipelines for action-conditioned world models and vision-language-action (VLA) models
- Develop and refine training methodologies, including fine-tuning, reinforcement learning, and large-scale multimodal learning
- Design and generate training and evaluation datasets from simulation, including environment setup, domain randomization, and sim-to-real transfer strategies
- Build distributed training infrastructure using PyTorch, FSDP, and DeepSpeed
- Work with multimodal data pipelines involving video, sensory inputs, and action sequences
- Evaluate model performance using both benchmark datasets and real-world deployment metrics
- Contributions research publications a plus
- Collaborate with industrial partners to adapt generative models for real-world physical AI applications
What the employer is looking for
- Experience with pre-training or post-training on large generative models (video, multimodal, or action-conditioned)
- Hands-on proficiency with PyTorch and distributed training frameworks (FSDP, DeepSpeed)
- Strong fundamentals in machine learning, optimization, and large-scale data processing
- Familiarity with VLMs, VLAs, or world models
- Background in robotics, embodied AI, or sim-to-real transfer is a plus
- Experience with video understanding or temporal reasoning is a plus
- BS/MS/PhD in Computer Science, Machine Learning, Robotics, or a related field