3 papers
cs.CV2026
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-t…
cs.CV2025
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Xiaohu Huang, Haoyang He, Hao Zhou +3
In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific archit…
cs.CV2025
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
Xiaohu Huang, Jingjing Wu, Qunyi Xie +1
Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of…