3 papers
cs.AI2026
Large Vision-Language Models Get Lost in Attention
Gongli Xi, Ye Tian, Mengyu Yang +5
Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer…
cs.CV2026
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
Gongli Xi, Kun Wang, Zeming Gao +4
Large multimodal reasoning models solve challenging visual problems via explicit long-chain inference: they gather visual clues from images and decode clues into textual tokens. Ye…
cs.CV2024
AdaViPro: Region-based Adaptive Visual Prompt for Large-Scale Models Adapting
Mengyu Yang, Ye Tian, Lanshan Zhang +3
Recently, prompt-based methods have emerged as a new alternative `parameter-efficient fine-tuning' paradigm, which only fine-tunes a small number of additional parameters while kee…