4 papers · 1 filter
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
Junha Song, Byeongho Heo, Geonmo Gu +3
When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast…
RL makes MLLMs see better than SFT
Junha Song, Sangdoo Yun, Dongyoon Han +2
A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarka…
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
Junha Song, Yongsik Jo, So Yeon Min +4
Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal la…
Is user feedback always informative? Retrieval Latent Defending for Semi-Supervised Domain Adaptation without Source Data
Junha Song, Tae Soo Kim, Junha Kim +3
This paper aims to adapt the source model to the target environment, leveraging small user feedback (i.e., labeled target data) readily available in real-world applications. We fin…