2 citations · 2 across the 5 of their papers we have counts for
9 papers
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
Chenyue Zhou, Jiayi Tuo, Shitong Qin +7
The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to s…
Geometric Prior-Guided Federated Prompt Calibration
Fei Luo, Ziwei Zhao, Mingxuan Wang +5
Federated Prompt Learning (FPL) offers a parameter-efficient solution for collaboratively training large models, but its performance is severely hindered by data heterogeneity, whi…
From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
Chenyue Zhou, Mingxuan Wang, Yanbiao Ma +19
Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incohere…
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
Yanbiao Ma, Wei Dai, Bowei Liu +5
Despite the fast progress of deep learning, one standing challenge is the gap of the observed training samples and the underlying true distribution. There are multiple reasons for…
Compositional Attribute Imbalance in Vision Datasets
Jiayi Chen, Yanbiao Ma, Andi Zhang +3
Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define…
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
Wenke Huang, Jian Liang, Xianda Guo +14
Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs dem…