4 papers · 1 filter
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
Wei Ao, Lan Wang, Vishnu Naresh Boddeti
The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language comp…
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Yujiang Pu, Zhanbo Huang, Vishnu Boddeti +1
Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image…
CryptoFace: End-to-End Encrypted Face Recognition
Wei Ao, Vishnu Naresh Boddeti
Face recognition is central to many authentication, security, and personalized applications. Yet, it suffers from significant privacy risks, particularly arising from unauthorized…
SEAL: Semantic Attention Learning for Long Video Representation
Lan Wang, Yujia Chen, Du Tran +2
Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must…