collaborators

6 papers

cs.CV2026

DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving

Xiaosong Jia, Yanhao Liu, Yu Hong +5

Feed-forward reconstruction has been progressed rapidly, with the Visual Geometry Grounded Transformer (VGGT) being a notable baseline. However, directly applying VGGT to autonomou…

cs.CV2026

VAR-3D: View-aware Auto-Regressive Model for Text-to-3D Generation via a 3D Tokenizer

Zongcheng Han, Dongyan Cao, Haoran Sun +1

Recent advances in auto-regressive transformers have achieved remarkable success in generative modeling. However, text-to-3D generation remains challenging, primarily due to bottle…

cs.GR2025

Topology-Aware Optimization of Gaussian Primitives for Human-Centric Volumetric Videos

Yuheng Jiang, Chengcheng Guo, Yize Wu +9

Volumetric video is emerging as a key medium for digitizing the dynamic physical world, creating the virtual environments with six degrees of freedom to deliver immersive user expe…

cs.GR2025

BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video

Yu Hong, Yize Wu, Zhehao Shen +5

Volumetric video enables immersive experiences by capturing dynamic 3D scenes, enabling diverse applications for virtual reality, education, and telepresence. However, traditional…

cs.CV2025

Towards Generalized and Training-Free Text-Guided Semantic Manipulation

Yu Hong, Xiao Cai, Pengpeng Zeng +4

Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addi…

cs.CV2025

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

Yuheng Jiang, Zhehao Shen, Chengcheng Guo +5

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiti…