From the 1 of 27 linked papers with an AI index.
27 papers
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Quanjian Song, Yiren Song, Kelly Peng +2
WorldWander is a framework that translates video content between first‑person (egocentric) and third‑person (exocentric) views using video diffusion transformers and in‑context lea…
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
Heyuan Gao, Bangxun Tang, Yiren Song +4
We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: generating dynamic backgrounds…
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
DataFlow Team, Bohan Zeng, Daili Hua +39
World models have garnered significant attention as a promising research direction in artificial intelligence, yet a clear and unified definition remains lacking. In this paper, we…
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
Yiren Song, Yihan Wang, Xiyao Deng +2
Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generatio…
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
Yiren Song, Huilin Zhong, Kevin Qinghong Lin +2
We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or actor replacement while strictly…
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
Yiren Song, Wangzi Yao, Haofan Wang +1
Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylizati…