27 citations · 28 across the 3 of their papers we have counts for
5 papers · 1 filter
LARE: Low-Attention Region Encoding for Text-Image Retrieval
Abdulmalik Alquwayfili, Faisal Almeshal, Jumanah Almajnouni +8
Image retrieval in crowded scenes is particularly challenging due to the salience bias of conventional visual encoders, which tend to focus on dominant objects while neglecting low…
Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search
Faisal Aljehrai, Mohammed A. Alkhrashi, Alreem Almuhrij +6
Video semantic search in densely crowded scenes remains a challenging task due to visual encoders tendency to prioritize salient foreground regions while neglecting contextually im…
VelocityNet: Real-Time Crowd Anomaly Detection via Person-Specific Velocity Analysis
Fatima AlGhamdi, Omar Alharbi, Abdullah Aldwyish +3
Detecting anomalies in crowded scenes is challenging due to severe inter-person occlusions and highly dynamic, context-dependent motion patterns. Existing approaches often struggle…
End-to-End Multimodal Representation Learning for Video Dialog
Huda Alamri, Anthony Bilic, Michael Hu +2
Video-based dialog task is a challenging multimodal learning task that has received increasing attention over the past few years with state-of-the-art obtaining new performance rec…
Audio-Visual Scene-Aware Dialog
Huda Alamri, Vincent Cartillier, Abhishek Das +9
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history…