5 papers
POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking
Zhangheng LI, Jianing Zhu, Junyuan Hong +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and visual data, where privacy-sen…
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
Stanislav Panev, Minhyek Jeon, Vaishnavi Khindkar +5
Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes. Beyond data augmentation, their potential as diagnostic t…
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
Yasiru Ranasinghe, Elim Schenck, Florence Yellin +3
Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge RGB-based semantics. We pres…
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
Xiao Fang, Minhyek Jeon, Zheyang Qin +5
Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state…
2D-3D Attention and Entropy for Pose Robust 2D Facial Recognition
J. Brennan Peace, Shuowen Hu, Benjamin S. Riggan
Despite recent advances in facial recognition, there remains a fundamental issue concerning degradations in performance due to substantial perspective (pose) differences between en…