2 papers
cs.CV2026
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
Wei Ao, Lan Wang, Vishnu Naresh Boddeti
The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language comp…
cs.CV2025
SEAL: Semantic Attention Learning for Long Video Representation
Lan Wang, Yujia Chen, Du Tran +2
Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must…