11 papers
TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization
Weiliang Chen, Yuanhui Huang, Xuebo Wang +1
Video tokenization is fundamental to scalable video generation, as the number of tokens directly determines the computational cost and the length of videos that can be modeled. Exi…
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
Weiliang Chen, Jiayi Bi, Yuanhui Huang +2
Generative models have shown great promise for novel view synthesis (NVS) by leveraging strong image generation priors. However, existing approaches typically follow a 2D inpaintin…
IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation
Yuqi Wu, Tianyu Hu, Wenzhao Zheng +4
Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundat…
Moaw: Unleashing Motion Awareness for Video Diffusion Models
Tianqi Zhang, Ziyi Wang, Wenzhao Zheng +5
Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks suc…
Terra: Explorable Native 3D World Model with Point Latents
Yuanhui Huang, Weiliang Chen, Wenzhao Zheng +4
World models have garnered increasing attention for comprehensive modeling of the real world. However, most existing methods still rely on pixel-aligned representations as the basi…
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
Yuqi Wu, Wenzhao Zheng, Sicheng Zuo +3
3D occupancy prediction provides a comprehensive description of the surrounding scenes and has become an essential task for 3D perception. Most existing methods focus on offline pe…