8 papers · 1 filter
Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models
Izaldein Al-Zyoud, Abdulmotaleb El Saddik
Frozen self-supervised vision models can align parts of generic objects, but it remains unclear whether this correspondence extends to human faces, where global layout is shared wh…
PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments
Minghao Zou, Qingtian Zeng, Shangkun Liu +5
Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets primarily fo…
Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating
Rui Yao, Yuhong Zhang, Kunyang Sun +4
Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for object tracking. However, ef…
Segmentation-Guided Spatial Indexing for Generalizable and Explainable Deepfake Detection
Izaldein Al-Zyoud, Abdulmotaleb El Saddik
We introduce segmentation-guided spatial indexing for generalizable and explainable deepfake detection. The key idea reverses the standard design order: rather than pooling all fac…
LOD-Net: Locality-Aware 3D Object Detection Using Multi-Scale Transformer Network
Mustaqeem Khan, Aidana Nurakhmetova, Wail Gueaieb +1
3D object detection in point cloud data remains a challenging task due to the sparsity and lack of global structure inherent in the input. In this work, we propose a novel Multi-Sc…
Real-Time Oriented Object Detection Transformer in Remote Sensing Images
Zeyu Ding, Yong Zhou, Jiaqi Zhao +4
Recent real-time detection transformers have gained popularity due to their simplicity and efficiency. However, these detectors do not explicitly model object rotation, especially…