From the 1 of 9 linked papers with an AI index.
4 papers · 1 filter
Interpretability Transfer from Language to Vision via Sparse Autoencoders
Alexey Kravets, Da Li, Chuan Li +2
Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambigui…
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Zixu Cheng, Da Li, Jian Hu +4
Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language…
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
Wenqing Wang, Da Li, Xiatian Zhu +1
Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adapt…
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
Lihao Liu, Yan Wang, Biao Yang +10
Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct applic…