10 papers · 1 filter
ActiveScope: Actively Seeking and Correcting Perception for MLLMs
Yajing Wang, Chao Bi, Junshu Sun +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, yet still struggle with fine-grained perception in high-resolution images. Whil…
Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation
Tiantian Dang, Chao Bi, Shufan Shen +3
Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and restricts broader practical deplo…
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
Shufan Shen, Zhaobo Qi, Junshu Sun +3
The visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new…
Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
Shufan Shen, Junshu Sun, Shuhui Wang +1
Parameter-efficient fine-tuning (PEFT) aims to adapt pre-trained vision models to downstream tasks. Among PEFT paradigms, sparse tuning achieves remarkable performance by adjusting…
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
Shufan Shen, Junshu Sun, Qingming Huang +1
The alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the a…
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
Yufan Zhou, Zhaobo Qi, Lingshuai Lin +5
In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual obser…