14 citations · 14 across the 3 of their papers we have counts for
4 papers
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
Zhaozhi Wang, Tong Zhang, Mingyue Guo +2
Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language alignment, yet they remain limited in visual-spatial reasoning. We first identify that…
MGPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
Mingshuang Luo, Ruibing Hou, Zhuo Li +4
This paper presents MGPT, an advanced ultimodal, ultitask framework for otion comprehension and generation. MGPT operates on three funda…
Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting
Mingyue Guo, Binghui Chen, Zhaoyi Yan +2
Multidomain crowd counting aims to learn a general model for multiple diverse datasets. However, deep networks prefer modeling distributions of the dominant domains instead of all…
Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
Mingyue Guo, Li Yuan, Zhaoyi Yan +3
Crowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios, however, regressors are challenged by uncontrol…