9 papers
ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
Zhexiao Xiong, Yizhi Song, Hao Kang +11
Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions corresp…
Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
Feng Qiao, Zhaochong An, Zhexiao Xiong +2
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the orig…
GenOpticalFlow: A Generative Approach to Unsupervised Optical Flow Learning
Yixuan Luo, Feng Qiao, Zhexiao Xiong +2
Optical flow estimation is a fundamental problem in computer vision, yet the reliance on expensive ground-truth annotations limits the scalability of supervised approaches. Althoug…
Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
Yiren Lu, Xin Ye, Burhaneddin Yaman +4
Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning fo…
PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment
Zhexiao Xiong, Yizhi Song, Liu He +4
Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing…
PanoDreamer: Consistent Text to 360-Degree Scene Generation
Zhexiao Xiong, Zhang Chen, Zhong Li +2
Automatically generating a complete 3D scene from a text description, a reference image, or both has significant applications in fields like virtual reality and gaming. However, cu…