2 papers
cs.SD2026
MATS: An Audio Language Model under Text-only Supervision
Wen Wang, Ruibing Hou, Hong Chang +2
Large audio-language models (LALMs), built upon powerful Large Language Models (LLMs), have exhibited remarkable audio comprehension and reasoning capabilities. However, the traini…
cs.CV2025
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
Keliang Li, Hongze Shen, Hao Shi +9
The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose…