7 papers
Plan-X: Instruct Video Generation via Semantic Planning
Lun Huang, You Xie, Hongyi Xu +7
Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This lim…
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
Yunlong Deng, Guangyi Chen, Tianpei Gu +4
Vision-Language Models (VLMs) integrate visual knowledge with the analytical capabilities of Large Language Models (LLMs) through supervised visual instruction tuning, using image-…
X-Streamer: Unified Human World Modeling with Audiovisual Interaction
You Xie, Tianpei Gu, Zenan Li +7
We introduce X-Streamer, an end-to-end multimodal human world modeling framework for building digital human agents capable of infinite interactions across text, speech, and video w…
Lynx: Towards High-Fidelity Personalized Video Generation
Shen Sang, Tiancheng Zhi, Tianpei Gu +2
We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introd…
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
Guoxian Song, Hongyi Xu, Xiaochen Zhao +5
We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike…
X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
Chenxu Zhang, Zenan Li, Hongyi Xu +8
We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input…