2 papers
cs.CV2025
Dynamic Token Reduction during Generation for Vision Language Models
Xiaoyu Liang, Chaofeng Guan, Jiaying Lu +3
Vision-Language Models (VLMs) have achieved notable success in multimodal tasks but face practical limitations due to the quadratic complexity of decoder attention mechanisms and a…
cs.CV2024
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
Xiaoyu Liang, Jiayuan Yu, Lianrui Mu +7
Although Visual-Language Models (VLMs) have shown impressive capabilities in tasks like visual question answering and image captioning, they still struggle with hallucinations. Ana…