3 papers
cs.CV2026
DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval
Kaleem Ullah, Altaf Hussain, Muhammad Munsif +1
In today's media-driven world, the exponential growth of video content across domains such as surveillance, education, and entertainment has made retrieving semantically relevant v…
cs.CV2025
AVAR-Net: A Lightweight Audio-Visual Anomaly Recognition Framework with a Benchmark Dataset
Amjid Ali, Zulfiqar Ahmad Khan, Altaf Hussain +3
Anomaly recognition plays a vital role in surveillance, transportation, healthcare, and public safety. However, most existing approaches rely solely on visual data, making them unr…
cs.CV2025
Pseudo-Labeling Driven Refinement of Benchmark Object Detection Datasets via Analysis of Learning Patterns
Min Je Kim, Muhammad Munsif, Altaf Hussain +2
Benchmark object detection (OD) datasets play a pivotal role in advancing computer vision applications such as autonomous driving, and surveillance, as well as in training and eval…