15 papers
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning
Jiacheng Yang, Tongying Xiao, Yunkai Dang +7
Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facili…
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Jiacheng Yang, Anqi Chen, Yunkai Dang +5
Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…
The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models
Liuyuan Wen, Xun Zhu, Lihao Huang +2
Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output. By analyzing the residual str…
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
Junyuan Ma, Xunzhi Xiang, Wenbin Li +2
Vision foundation models (VFMs) have achieved strong performance across various vision tasks. However, it still remains challenging to apply VFMs for cross-domain few-shot segmenta…
Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models
Yunkai Dang, Yifan Jiang, Yizhu Jiang +3
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in p…
Understanding and Enforcing Weight Disentanglement in Task Arithmetic
Shangge Liu, Yuehan Yin, Lei Wang +5
Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…