3 papers
cs.LG2026
Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
Hahyeon Choi, Nojun Kwak
We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fi…
cs.CV2025
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
Hahyeon Choi, Junhoo Lee, Nojun Kwak
Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to…
cs.CL2024
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
Dayoon Ko, Jinyoung Kim, Hahyeon Choi +1
In the real world, knowledge is constantly evolving, which can render existing knowledge-based datasets outdated. This unreliability highlights the critical need for continuous upd…