1 citations · 2 across the 46 of their papers we have counts for
34 papers
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
Bojia Zi, Xiaoyan Yang, Yu Zhou +7
Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, t…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Yuehao Huang, Yunzi Wu, Xiaotao Zhang +7
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observati…
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
Dongxu Ge, Shansong Liu, Cheng Gong +3
As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
Yuyang Huang, Yabo Chen, Wenrui Dai +6
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters a…
ShotPlan: Cinematic Video Generation with Learnable Planning Token
Su Guo, Guangce Liu, Haosen Yang +7
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective mult…
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Yuanzhi Liang, Xufeng Zhan, Haibin Huang +2
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have adva…