activity
20242026
collaborators

8 papers

cs.SD2026

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

Hyebin Cho, Suho Yoo, Jaehyuk Jang +2

While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, ca…

cs.SD2026

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

Hyebin Cho, Jaehyuk Jang, Changick Kim +1

Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance…

cs.MM2026

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

Inyong Koo, yeeun Seong, Minseok Son +2

Audio-visual emotion recognition (AVER) methods typically fuse utterance-level features, and even frame-level attention models seldom address the frame-rate mismatch across modalit…

cs.CV2025

Towards Efficient Vision State Space Models via Token Merging

Jinyoung Park, Minseok Son, Changick Kim

State Space Models (SSMs) have emerged as powerful architectures in computer vision, yet improving their computational efficiency remains crucial for practical and scalable deploym…

cs.CV2025

Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Sangmin Woo, Donguk Kim, Jaehyuk Jang +2

Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or mislead…

cs.CV2024

Difficulty-aware Balancing Margin Loss for Long-tailed Recognition

Minseok Son, Inyong Koo, Jinyoung Park +1

When trained with severely imbalanced data, deep neural networks often struggle to accurately recognize classes with only a few samples. Previous studies in long-tailed recognition…