most citedPursuing Better Decision Boundaries for Long-Tailed Object Detection via Category Information Amount

2 citations · 2 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CV2026

MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers

Chenyue Zhou, Jiayi Tuo, Shitong Qin +7

The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to s…

cs.LG2025

Geometric Prior-Guided Federated Prompt Calibration

Fei Luo, Ziwei Zhao, Mingxuan Wang +5

Federated Prompt Learning (FPL) offers a parameter-efficient solution for collaboratively training large models, but its performance is severely hindered by data heterogeneity, whi…

cs.AI2025

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma +19

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incohere…

cs.CV2025

Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency

Yanbiao Ma, Wei Dai, Bowei Liu +5

Despite the fast progress of deep learning, one standing challenge is the gap of the observed training samples and the underlying true distribution. There are multiple reasons for…

cs.CV2025

Compositional Attribute Imbalance in Vision Datasets

Jiayi Chen, Yanbiao Ma, Andi Zhang +3

Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define…

cs.CL2025

Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

Wenke Huang, Jian Liang, Xianda Guo +14

Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs dem…