2 papers
cs.AI2025
From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
Chenyue Zhou, Mingxuan Wang, Yanbiao Ma +19
Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incohere…
cs.CV2025
Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
Yanbiao Ma, Wei Dai, Bowei Liu +5
Despite the fast progress of deep learning, one standing challenge is the gap of the observed training samples and the underlying true distribution. There are multiple reasons for…