4 papers
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
Jing Zuo, Lingzhou Mu, Fan Jiang +3
Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to jointly understand multimodal instructions and visual-spatial context while…
Bridging the gap between training and inference in LM-based TTS models
Ruonan Zhang, Lingzhou Mu, Xixin Wu +1
Recent advancements in text-to-speech (TTS) have shown that language model (LM) based systems offer competitive performance compared to traditional approaches. However, in training…
FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
Lingzhou Mu, Qiang Wang, Fan Jiang +4
Human-Scene Interaction (HSI) seeks to generate realistic human behaviors within complex environments, yet it faces significant challenges in handling long-horizon, high-level task…
FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model
Lingzhou Mu, Baiji Liu, Ruonan Zhang +4
Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressio…