Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
Wei Ao, Lan Wang, Vishnu Naresh Boddeti
The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language comp…
cs.CV2024
SEAL: Semantic Attention Learning for Long Video Representation
Lan Wang, Yujia Chen, Du Tran +2
Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must…
cs.CV2024
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
Lan Wang, Vishnu Boddeti, Sernam Lim
We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method…