4 citations · 6 across the 12 of their papers we have counts for
7 papers · 1 filter
Human detectors are surprisingly powerful reward models
Kumar Ashutosh, XuDong Wang, Xi Yin +4
Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
Bolin Lai, Xudong Wang, Saketh Rambhatla +4
Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-c…
Generating Multi-Image Synthetic Data for Text-to-Image Customization
Nupur Kumari, Xi Yin, Jun-Yan Zhu +2
Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive…
Diffusion Autoencoders are Scalable Image Tokenizers
Yinbo Chen, Rohit Girdhar, Xiaolong Wang +2
Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models. We present a simple diffusion tokenizer (DiTo) t…
LLMs can see and hear without any training
Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2
We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ab…
CAT: Content-Adaptive Image Tokenization
Junhong Shen, Kushal Tirumala, Michihiro Yasunaga +4
Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity. To address this, we introduce Conte…