5 papers
Image Diffusion Preview with Consistency Solver
Fu-Yun Wang, Hao Zhou, Liangzhe Yuan +8
The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employin…
Neptune: The Long Orbit to Benchmarking Long Video Understanding
Arsha Nagrani, Mingda Zhang, Ramin Mehran +10
We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…
Extending Video Masked Autoencoders to 128 frames
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal +8
Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked…
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
Lichang Chen, Hexiang Hu, Mingda Zhang +8
We introduce OmnixR, an evaluation suite designed to benchmark SoTA Omni-modality Language Models, such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities s…
Epsilon-VAE: Denoising as Visual Decoding
Long Zhao, Sanghyun Woo, Ziyu Wan +6
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data,…