AIMovieJobs.com
Login/Signup

AI Video Generation

Research Scientist, Foundation Model (Video Generation)

Pika

External listing
Palo Alto HQOnsiteFull-time$185,000–$400,000/year
External listing from the employer

Originally published by Pika through its Ashby careers source. Applications go to the employer's website, and AIMovieJobs is not the hiring employer.

View original listing

Quick facts

Palo Alto HQLocationOnsiteWorkplaceFull-timeJob type$185,000–$400,000/yearCompensationAI Video GenerationCategoryResearchDepartmentMay 16, 2026PostedAug 24, 2026Last verifiedFeb 20, 2027Expires

Role overview

As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal pre-training/mid-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale. What You’ll Do Lead research and development on pre-training and mid-training of multimodal foundation models at scale. Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.

What you'll do

  • Lead research and development on pre-training and mid-training of multimodal foundation models at scale.
  • Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.
  • Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.
  • Advance state-of-the-art techniques in diffusion, autoregressive, and other generative models for large-scale pre-training and fine-tuning.
  • Identify, create, and leverage large, high-quality cross-modal datasets.
  • Bring research advancements into production-ready systems in collaboration with engineering and product teams.
  • Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.
  • Stay at the forefront of foundational model and real-time multimodal AI research.

What the employer is looking for

  • 5+ years of research experience in large-scale pre-training/mid-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.
  • Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).
  • Extensive hands-on experience with large-scale multimodal model design, training, and deployment.
  • Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).
  • Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.
  • Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.
  • Excellent communication and collaboration skills, and a passion for building creative enabling technology.