activity
20172026
most citedAttentional Pooling for Action Recognition

209 citations · 264 across the 23 of their papers we have counts for

collaborators
Showing cs.CVShow all

31 papers · 1 filter

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…

cs.CV2025

Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective

Bolin Lai, Xudong Wang, Saketh Rambhatla +4

Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-c…

cs.CV2025

Diffusion Autoencoders are Scalable Image Tokenizers

Yinbo Chen, Rohit Girdhar, Xiaolong Wang +2

Tokenizing images into compact visual representations is a key step in learning efficient and high-quality image generative models. We present a simple diffusion tokenizer (DiTo) t…

cs.CV2025

LLMs can see and hear without any training

Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2

We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ab…

cs.CV2024

MotiF: Making Text Count in Image Animation with Motion Focal Loss

Shijie Wang, Samaneh Azadi, Rohit Girdhar +3

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing m…

cs.CV2024

Human Action Anticipation: A Survey

Bolin Lai, Sam Toyer, Tushar Nagarajan +7

Predicting future human behavior is an increasingly popular topic in computer vision, driven by the interest in applications such as autonomous vehicles, digital assistants and hum…