3 papers
cs.AI2025
From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
Chenyue Zhou, Mingxuan Wang, Yanbiao Ma +19
Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incohere…
cs.CV2025
Geometric Origins of Bias in Deep Neural Networks: A Human Visual System Perspective
Yanbiao Ma, Bowei Liu, Andi Zhang
Bias formation in deep neural networks (DNNs) remains a critical yet poorly understood challenge, influencing both fairness and reliability in artificial intelligence systems. Insp…
cs.CV2025
Compositional Attribute Imbalance in Vision Datasets
Jiayi Chen, Yanbiao Ma, Andi Zhang +3
Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define…