9 papers
Natural Language Camera Movement Understanding
Yuwen Tan, Joey Huang, Jin Huang +2
Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing v…
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies…
Image Diffusion Preview with Consistency Solver
Fu-Yun Wang, Hao Zhou, Liangzhe Yuan +8
The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employin…
VideoPrism: A Foundational Visual Encoder for Video Understanding
Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus…
Epsilon-VAE: Denoising as Visual Decoding
Long Zhao, Sanghyun Woo, Ziyu Wan +6
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…
Neptune: The Long Orbit to Benchmarking Long Video Understanding
Arsha Nagrani, Mingda Zhang, Ramin Mehran +10
We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…