From the 1 of 9 linked papers with an AI index.
9 papers
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5
The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…
SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval
Yuxiao Luo, Da Li, Mingjie Zhang +3
LLM-based retrievers have become a fundamental component of modern information retrieval systems. The paradigm of "rewrite-then-retriev" introduces explicit reasoning before retrie…
Interpretability Transfer from Language to Vision via Sparse Autoencoders
Alexey Kravets, Da Li, Chuan Li +2
Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambigui…
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Zixu Cheng, Da Li, Jian Hu +4
Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language…
Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning
Zechang Xiong, Da Li, Kexin Tang +3
Multimodal models often converge to a dominant-modality solution, in which a stronger, faster-converging modality overshadows weaker ones. This modality imbalance causes suboptimal…
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
Wenqing Wang, Da Li, Xiatian Zhu +1
Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adapt…