activity
20232026
most citedOn Bringing Robots Home

4 citations · 6 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…

cs.CV2025

Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective

Bolin Lai, Xudong Wang, Saketh Rambhatla +4

Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-c…

cs.CV2025

Generating Multi-Image Synthetic Data for Text-to-Image Customization

Nupur Kumari, Xi Yin, Jun-Yan Zhu +2

Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive…

cs.CV2025

Diffusion Autoencoders are Scalable Image Tokenizers

Yinbo Chen, Rohit Girdhar, Xiaolong Wang +2

Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models. We present a simple diffusion tokenizer (DiTo) t…

cs.CV2025

LLMs can see and hear without any training

Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2

We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ab…

cs.CV2025

CAT: Content-Adaptive Image Tokenization

Junhong Shen, Kushal Tirumala, Michihiro Yasunaga +4

Most existing image tokenizers encode images into a fixed number of tokens or patches, overlooking the inherent variability in image complexity. To address this, we introduce Conte…