15 papers
Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention
Jiaqian Li, Yanshu Li, Ligong Han +2
Implicit in-context learning (ICL) has newly emerged as a promising paradigm that simulates ICL behaviors in the representation space of large language models (LLMs), aiming to att…
Personalize Your Large Vision-language Models With In-context Prompt Tuning
Yanshu Li, Jiaqian Li, Kuai Yu +4
Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This trend has driven growing inter…
Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
Yuefei Chen, Jiang Liu, Xiaodong Lin +1
Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reasoning, and instruction-following tasks…
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
Liwei Che, Zhiyu Xue, Yihao Quan +7
Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual object and then add them all up…
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
Yanchuan Tang, Taowen Wang, Yuefei Chen +3
Vision-Language-Action (VLA) models enable general-purpose robotic policies by mapping visual observations and language instructions to low-level actions, but they often lack relia…
Improving Visual Reasoning with Iterative Evidence Refinement
Zeru Shi, Kai Mei, Yihao Quan +2
Vision language models (VLMs) are increasingly capable of reasoning over images, but robust visual reasoning often requires re-grounding intermediate steps in the underlying visual…