8 papers
SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation
Zihao Zhang, Haoyu Zhao, Siqian Yang +3
Despite the success of diffusion models in Video Frame Interpolation (VFI), existing methods still suffer from two critical limitations. First, latent diffusion inevitably loses fi…
PreferThinker: Reasoning-based Personalized Image Preference Assessment
Shengqi Xu, Xinpeng Zhou, Yabo Zhang +6
Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing m…
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
Zhihao He, Tieyuan Chen, Kangyu Wang +6
Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
Haoyu Zhao, Jiaxi Gu, Shicong Wang +4
The explosive growth of video streaming presents challenges in achieving high accuracy and low training costs for video-language retrieval. However, existing methods rely on large-…
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
Quan-Sheng Zeng, Yunheng Li, Qilong Wang +4
Visual token compression is critical for Large Vision-Language Models (LVLMs) to efficiently process high-resolution inputs. Existing methods that typically adopt fixed compression…
Aligning Anime Video Generation with Human Feedback
Bingwen Zhu, Yudong Jiang, Baohan Xu +5
Anime video generation faces significant challenges due to the scarcity of anime data and unusual motion patterns, leading to issues such as motion distortion and flickering artifa…