7 papers
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
Ling Lin, Yang Bai, Congcong Zhu +6
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning pro…
See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs
Yuqing Lei, Wenbo Lyu, Yingjun Du +3
Large Vision-Language Models (LVLMs) excel at multimodal tasks but remain prone to object hallucinations. Prior training-free remedies often uniformly strengthen visual signals, wh…
Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning
Zhirui Chen, Ziwei Chen, Ling Shao
Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by finite context windows and the…
StructKV: Preserving the Structural Skeleton for Scalable Long-Context Inference
Zhirui Chen, Peiyang Liu, Ling Shao
As Large Language Models (LLMs) scale to support context windows exceeding one million tokens, the linear growth of Key-Value (KV) cache imposes severe memory capacity and bandwidt…
Self-Consolidation for Self-Evolving Agents
Hongzhuo Yu, Fei Zhu, Guo-Sen Xie +1
While large language model (LLM) agents have demonstrated impressive problem-solving capabilities, they typically operate as static systems, lacking the ability to evolve through l…
MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
Yuqing Lei, Yingjun Du, Yawen Huang +2
Vision-language models (VLMs) such as CLIP exhibit strong zero-shot generalization but remain sensitive to domain shifts at test time. Test-time prompt tuning (TPT) mitigates this…