5 papers · 1 filter
ActiveScope: Actively Seeking and Correcting Perception for MLLMs
Yajing Wang, Chao Bi, Junshu Sun +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, yet still struggle with fine-grained perception in high-resolution images. Whil…
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
Shufan Shen, Zhaobo Qi, Junshu Sun +3
The visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new…
Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
Shufan Shen, Junshu Sun, Shuhui Wang +1
Parameter-efficient fine-tuning (PEFT) aims to adapt pre-trained vision models to downstream tasks. Among PEFT paradigms, sparse tuning achieves remarkable performance by adjusting…
VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
Shufan Shen, Junshu Sun, Qingming Huang +1
The alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the a…
Expanding Sparse Tuning for Low Memory Usage
Shufan Shen, Junshu Sun, Xiangyang Ji +2
Parameter-efficient fine-tuning (PEFT) is an effective method for adapting pre-trained vision models to downstream tasks by tuning a small subset of parameters. Among PEFT methods,…