6 papers
PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning
Chen Yang, Shenxiang Zeng, Haoyang Zhao +6
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to int…
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
Ruina Hu, Chen Wang, Lai Wei +5
Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only re…
Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces
Chen Yang, Guanxin Lin, Youquan He +10
Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a single image may admit multiple pl…
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
Jun Zhan, Mingyang Han, Yuxuan Xie +11
Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress h…
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
Jihao Gu, Yingyao Wang, Pi Bu +17
The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge…
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
Liangyu Xu, Yingxiu Zhao, Jingyun Wang +9
Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large l…