activity
20242026
collaborators

9 papers

cs.CV2026

Natural Language Camera Movement Understanding

Yuwen Tan, Joey Huang, Jin Huang +2

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing v…

cs.CV2026

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5

The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies…

cs.LG2026

Image Diffusion Preview with Consistency Solver

Fu-Yun Wang, Hao Zhou, Liangzhe Yuan +8

The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employin…

cs.CV2025

VideoPrism: A Foundational Visual Encoder for Video Understanding

Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…

cs.CV2025

Epsilon-VAE: Denoising as Visual Decoding

Long Zhao, Sanghyun Woo, Ziyu Wan +6

In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…

cs.LG2025

Neptune: The Long Orbit to Benchmarking Long Video Understanding

Arsha Nagrani, Mingda Zhang, Ramin Mehran +10

We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…