3d understanding 1multimodal large language models 1prior fusion 1spatial reasoning 1visual priors 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
The paper introduces ViPS, a framework that combines multiple visual priors from diverse foundation models using an Efficient Prior Proxy and Dynamic Prior Fusion to improve spatia…
cs.CV2025
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Xiaohu Huang, Hao Zhou, Haoyang He +3
In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific archit…
cs.CV2025
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
Xiaohu Huang, Jingjing Wu, Qunyi Xie +1
Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of…