RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
arXiv:2307.15988 · doi:10.1109/ACCESS.2023.3312017
Abstract
We present RGB-D-Fusion, a multi-modal conditional denoising diffusion probabilistic model to generate high resolution depth maps from low-resolution monocular RGB images of humanoid subjects. RGB-D-Fusion first generates a low-resolution depth map using an image conditioned denoising diffusion probabilistic model and then upsamples the depth map using a second denoising diffusion probabilistic model conditioned on a low-resolution RGB-D image. We further introduce a novel augmentation technique, depth noise augmentation, to increase the robustness of our super-resolution model.
References in corpus (16)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Generative Adversarial Networks
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Score-Based Generative Modeling through Stochastic Differential Equations
- Linformer: Self-Attention with Linear Complexity
- Classifier-Free Diffusion Guidance
- DreamFusion: Text-to-3D using 2D Diffusion
- ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
- Tackling the Generative Learning Trilemma with Denoising Diffusion GANs
- Monocular Depth Estimation: A Survey
- Novel View Synthesis with Diffusion Models
- LiDAR-guided Stereo Matching with a Spatial Consistency Constraint
- Monocular Depth Estimation using Diffusion Models
- DAG: Depth-Aware Guidance with Denoising Diffusion Probabilistic Models