7 papers
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
Chenggong Hu, Shaoyin Ma, Yi Wang +3
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability an…
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
Chenggong Hu, Yi Wang, Mengqi Xue +3
Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated T…
-RSMDE: 40 Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
Ruizhi Wang, Weihan Li, Zunlei Feng +5
Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy a…
DriveFix: Spatio-Temporally Coherent Driving Scene Restoration
Heyu Si, Brandon James Denis, Muyang Sun +9
Recent advancements in 4D scene reconstruction, particularly those leveraging diffusion priors, have shown promise for novel view synthesis in autonomous driving. However, these me…
From Rays to Projections: Better Inputs for Feed-Forward View Synthesis
Zirui Wu, Zeren Jiang, Martin R. Oswald +1
Feed-forward view synthesis models predict a novel view in a single pass with minimal 3D inductive bias. Existing works encode cameras as Plücker ray maps, which tie predictions t…
ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos
Zetong Zhang, Manuel Kaufmann, Lixin Xue +2
Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural…