collaborators

13 papers

cs.CV2026

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

Chuangxin Zhao, Canran Xiao, Siyuan Ma +5

Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…

cs.CV2026

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning

Hengbo Xu, Shengjie Jin, Yanbiao Ma +1

With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual…

cs.AI2026

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

Zhe Qian, Nianbing Su, Zhonghua Wang +6

Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self…

cs.AI2026

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models

Zhe Qian, Yanbiao Ma, Zhuohan Ouyang +7

Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucin…

cs.CV2026

MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers

Chenyue Zhou, Jiayi Tuo, Shitong Qin +7

The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to s…

cs.LG2025

Geometric Prior-Guided Federated Prompt Calibration

Fei Luo, Ziwei Zhao, Mingxuan Wang +5

Federated Prompt Learning (FPL) offers a parameter-efficient solution for collaboratively training large models, but its performance is severely hindered by data heterogeneity, whi…