7 papers
Tracking-by-detection in Multi-object Tracking: Survey and Experiments
Yujin Yang, Kyujin Shim, Kangwook Ko +1
Multi-object tracking (MOT) is an essential computer vision task that simultaneously tracks multiple objects in video sequences, with various applications in surveillance, autonomo…
Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models
Kangwook Ko, Jaehyuk Jang, Wonjun Lee +2
Removing a specific individual's information from multimodal large language models (MLLMs) is often needed after deployment, but existing methods rely on a retain set, which is har…
AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning
Wonjun Lee, Jaehyuk Jang, Kangwook Ko +2
Multimodal large language models (MLLMs) can memorize identity-specific facts about people in their fine-tuning data, creating privacy risks when a person requests deletion. Existi…
T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models
Jaehyuk Jang, Minseok Seo, Seungju Cho +2
Vision-language models (VLMs) achieve strong zero-shot recognition, but they remain highly vulnerable to adversarial perturbations. Recent test-time adaptations improve robustness…
SubT: Subspace Tuning for Few-shot Generalization of Audio-Language Models
Jaehyuk Jang, Kangwook Ko, Wonjun Lee +1
Few-shot parameter-efficient adaptation of pretrained Audio--Language Models (ALMs) often improves seen-class performance at the cost of unseen-class generalization, leading to the…
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
Jaehyuk Jang, Wonjun Lee, Kangwook Ko +1
Prompt tuning has achieved remarkable progress in vision-language models (VLMs) and is recently being adopted for audio-language models (ALMs). However, its generalization ability…