5 papers
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
Chong Gao, Jie Ma, Zhan Peng +5
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to shar…
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
Zhida Zhang, Jie Ma, Zhan Peng +5
The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantl…
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
Haofeng Liu, Yang Zhou, Ziheng Wang +6
Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors…
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
Jie Ma, Yihang Liu, Zhike Qiu +2
Are low-attention visual tokens truly redundant in vision-language reasoning? Existing pruning methods often assume so, ranking visual tokens by shallow text-to-image attention and…
MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation
Dongxia Liu, Jie Ma, Xiaochen Yang +7
The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-intensive and computationally expen…