AIMovieJobs.com
Login/Signup

AI Video Generation

ML Researcher - Image / Video Diffusion

Krea

External listing
San FranciscoOnsiteFull-timeCompensation not provided
External listing from the employer

Originally published by Krea through its Ashby careers source. Applications go to the employer's website, and AIMovieJobs is not the hiring employer.

View original listing

Quick facts

San FranciscoLocationOnsiteWorkplaceFull-timeJob typeCompensation not providedCompensationAI Video GenerationCategoryResearchDepartmentSep 1, 2026PostedSep 8, 2026Last verifiedMar 7, 2027Expires

Role overview

We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale. What you'll doTrain diffusion models for image and video generation on large GPU clusters. Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design. Experience in profiling and debugging large distributed training.

What you'll do

  • Train diffusion models for image and video generation on large GPU clusters.
  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.
  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.
  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

What the employer is looking for

  • Proven track record in working with image or video models at scale (publications or open-source contributions a plus).
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
  • Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.
  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
  • Being comfortable working in a goal-oriented research environment.
  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
  • Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.
  • Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.
  • Ability to iterate rapidly, and propose creative research directions.
  • Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.