collaborators

5 papers

cs.CL2026

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning

Shan Yang

We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures vision-language reasoning: t…

cs.CV2026

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models

Mingzhe Huang, Weijun Wang, Xin Ding +7

In Vision-Language Models (VLMs), processing a massive number of visual tokens incurs prohibitive computational overhead. While recent training-aware pruning methods attempt to sel…

eess.AS2026

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

Guanrou Yang, Tian Tan, Qian Chen +12

Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks current…

cs.MM2025

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Le Xu, Chenxing Li, Yong Ren +5

Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge t…

cs.MM2025

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu +7

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…