13 papers
Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs
Chuangxin Zhao, Canran Xiao, Siyuan Ma +5
Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
Hengbo Xu, Shengjie Jin, Yanbiao Ma +1
With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual…
SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
Zhe Qian, Nianbing Su, Zhonghua Wang +6
Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self…
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models
Zhe Qian, Yanbiao Ma, Zhuohan Ouyang +7
Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucin…
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
Chenyue Zhou, Jiayi Tuo, Shitong Qin +7
The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to s…
Geometric Prior-Guided Federated Prompt Calibration
Fei Luo, Ziwei Zhao, Mingxuan Wang +5
Federated Prompt Learning (FPL) offers a parameter-efficient solution for collaboratively training large models, but its performance is severely hindered by data heterogeneity, whi…