4 papers
Unified Multimodal Models as Auto-Encoders
Zhiyuan Yan, Kaiqing Lin, Zongjian Li +10
Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection,…
Surfel-based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction from Monocular Video
Yiqun Zhao, Chenming Wu, Binbin Huang +4
Efficient and accurate reconstruction of a relightable, dynamic clothed human avatar from a monocular video is crucial for the entertainment industry. This paper presents SGIA (Sur…
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
Hao Li, Chenming Wu, Ming Yuan +7
Comprehensive testing of autonomous systems through simulation is essential to ensure the safety of autonomous driving vehicles. This requires the generation of safety-critical sce…
DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
Hao Li, Yuanyuan Gao, Haosong Peng +7
Novel-view synthesis (NVS) approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, mak…