3 papers
cs.CL2026
CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage
Zhangwei Cao, Shuhan Fan, Yuting Wei +5
Evaluating Multimodal Large Language Models (MLLMs) in Chinese Cultural Heritage (CCH) requires fine-grained reasoning over visual, textual, stylistic, and historical clues. Howeve…
cs.CV2025
MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
Yuandong Wang, Yao Cui, Yuxin Zhao +3
Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reaso…
cs.RO2025
Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization
Jian Deng, Yuandong Wang, Yangfu Zhu +3
Robotic manipulation systems are increasingly deployed across diverse domains. Yet existing multi-modal learning frameworks lack inherent guarantees of geometric consistency, strug…