1 paper
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…