3 papers
cs.CV2026
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning
Jiacheng Yang, Tongying Xiao, Yunkai Dang +7
Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facili…
cs.CV2026
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Jiacheng Yang, Anqi Chen, Yunkai Dang +5
Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…
cs.AI2026
Understanding and Enforcing Weight Disentanglement in Task Arithmetic
Shangge Liu, Yuehan Yin, Lei Wang +5
Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…