11 papers
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Yichen Guo, Kai Tang, Fenglai Lin +5
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Kai Tang, Jinhao You, Bohua Zhang +6
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain su…
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
Junpeng Ma, Sashuai Zhou, Guanghao Li +9
Video Large Language Models (VLMs) have achieved remarkable success in video understanding, but the significant computational cost from processing dense frames severely limits thei…
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
Ruichuan An, Sihan Yang, Renrui Zhang +10
Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies have investigat…
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
Gaole Dai, Menghang Dong, Rongyu Zhang +3
The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the fun…
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8
In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To…