6 papers
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
Xiaosong Jia, Yanhao Liu, Yu Hong +5
Feed-forward reconstruction has been progressed rapidly, with the Visual Geometry Grounded Transformer (VGGT) being a notable baseline. However, directly applying VGGT to autonomou…
VAR-3D: View-aware Auto-Regressive Model for Text-to-3D Generation via a 3D Tokenizer
Zongcheng Han, Dongyan Cao, Haoran Sun +1
Recent advances in auto-regressive transformers have achieved remarkable success in generative modeling. However, text-to-3D generation remains challenging, primarily due to bottle…
Topology-Aware Optimization of Gaussian Primitives for Human-Centric Volumetric Videos
Yuheng Jiang, Chengcheng Guo, Yize Wu +9
Volumetric video is emerging as a key medium for digitizing the dynamic physical world, creating the virtual environments with six degrees of freedom to deliver immersive user expe…
BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video
Yu Hong, Yize Wu, Zhehao Shen +5
Volumetric video enables immersive experiences by capturing dynamic 3D scenes, enabling diverse applications for virtual reality, education, and telepresence. However, traditional…
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
Yu Hong, Xiao Cai, Pengpeng Zeng +4
Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addi…
RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance
Yuheng Jiang, Zhehao Shen, Chengcheng Guo +5
Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiti…