2 citations · 2 across the 14 of their papers we have counts for
10 papers · 1 filter
Information Density Imbalance in Visual Object Detection
Ziwei Zhao, Yanxi Lu, Yuwei Hu +8
In object detection, the number of instances is typically used to determine whether a dataset exhibits a long-tailed distribution, implicitly assuming that the model will perform p…
Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs
Chuangxin Zhao, Canran Xiao, Siyuan Ma +5
Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
Hengbo Xu, Shengjie Jin, Yanbiao Ma +1
With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual…
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
Chenyue Zhou, Jiayi Tuo, Shitong Qin +7
The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to s…
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
Yanbiao Ma, Wei Dai, Bowei Liu +5
Despite the fast progress of deep learning, one standing challenge is the gap of the observed training samples and the underlying true distribution. There are multiple reasons for…
Compositional Attribute Imbalance in Vision Datasets
Jiayi Chen, Yanbiao Ma, Andi Zhang +3
Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define…