10 papers · 1 filter
Human detectors are surprisingly powerful reward models
Kumar Ashutosh, XuDong Wang, Xi Yin +4
Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
Bolin Lai, Xudong Wang, Saketh Rambhatla +4
Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-c…
Generating Multi-Image Synthetic Data for Text-to-Image Customization
Nupur Kumari, Xi Yin, Jun-Yan Zhu +2
Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive…
Movie Gen: A Cast of Media Foundation Models
Adam Polyak, Amit Zohar, Andrew Brown +85
We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabili…
Diffusion Autoencoders are Scalable Image Tokenizers
Yinbo Chen, Rohit Girdhar, Xiaolong Wang +2
Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models. We present a simple diffusion tokenizer (DiTo) t…
LLMs can see and hear without any training
Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2
We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ab…