5 papers
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
Zheyu Zhang, Ziqi Pang, Shixing Chen +3
Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens…
RefTok: Reference-Based Tokenization for Video Generation
Xiang Fan, Xiaohang Sun, Kushan Thakkar +4
Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectivel…
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
Hanglei Zhang, Yiwei Guo, Zhihan Li +3
Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently h…
NowYouSee Me: Context-Aware Automatic Audio Description
Seon-Ho Lee, Jue Wang, David Fan +5
Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable int…
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning
Yicheng Wang, Zhikang Zhang, Jue Wang +6
In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from tw…