11 papers
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
Liyang Li, Wen Wang, Canyu Zhao +4
Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, ex…
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
Inclusion AI, Tiwei Bie, Haoxing Chen +15
We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its…
Exploring Spatial Intelligence from a Generative Perspective
Muzhi Zhu, Shunyao Jiang, Huanyi Zheng +9
Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern gener…
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
Canyu Zhao, Mingyu Liu, Wen Wang +5
Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narrati…
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
Yongtao Ge, Wenjia Wang, Yongfan Chen +4
Despite remarkable progress having been made on the problem of 3D human pose and shape estimation (HPS), current state-of-the-art methods rely heavily on either confined indoor moc…
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
Canyu Zhao, Yanlong Sun, Mingyu Liu +6
This paper's primary objective is to develop a robust generalist perception model capable of addressing multiple tasks under constraints of computational resources and limited trai…