5 papers
Feature-level Interaction Explanations in Multimodal Transformers
Yeji Kim, Housam Khalifa Bashier Babiker, Mi-Young Kim +1
Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision. Most existing multimodal explainable AI (MXAI) methods ext…
A More Word-like Image Tokenization for MLLMs
Hyun Lee, Hyemin Jeong, Yejin Kim +4
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding…
Enhancing Multi-Image Understanding through Delimiter Token Scaling
Minyoung Lee, Yeji Park, Dongjun Hwang +3
Large Vision-Language Models (LVLMs) achieve strong performance on single-image tasks, but their performance declines when multiple images are provided as input. One major reason i…
Knowledge Vector Weakening: Efficient Training-free Unlearning for Large Vision-Language Models
Yejin Kim, Dongjun Hwang, Sungmin Cha +1
Large Vision-Language Models (LVLMs) are widely adopted for their strong multimodal capabilities, yet they raise serious concerns such as privacy leakage and harmful content genera…
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
Yejin Kim, Eunwon Kim, Buru Chang +1
LLMs have demonstrated remarkable performance across various tasks but face challenges related to unintentionally generating outputs containing sensitive information. A straightfor…