works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.CV2026

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

Xiao Lin, Xiaohu Huang, Kai Han

The paper introduces ViPS, a framework that combines multiple visual priors from diverse foundation models using an Efficient Prior Proxy and Dynamic Prior Fusion to improve spatia…

cs.CV2026

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Donghao Zhou, Guisheng Liu, Hao Yang +9

In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference image…

cs.CV2025

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Xiaohu Huang, Hao Zhou, Haoyang He +3

In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific archit…

cs.CV2025

3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding

Xiaohu Huang, Jingjing Wu, Qunyi Xie +1

Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of…

cs.CV2024

PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Xiaohu Huang, Hao Zhou, Kai Han

In this paper, we introduce PruneVid, a visual token pruning method designed to enhance the efficiency of multi-modal video understanding. Large Language Models (LLMs) have shown p…