RGBX: Image decomposition and synthesis using material- and lighting-aware diffusion models
arXiv:2405.00666 · doi:10.1145/3641519.3657445
Abstract
The three areas of realistic forward rendering, per-pixel inverse rendering, and generative image synthesis may seem like separate and unrelated sub-fields of graphics and vision. However, recent work has demonstrated improved estimation of per-pixel intrinsic channels (albedo, roughness, metallicity) based on a diffusion architecture; we call this the RGBX problem. We further show that the reverse problem of synthesizing realistic images given intrinsic channels, XRGB, can also be addressed in a diffusion framework. Focusing on the image domain of interior scenes, we introduce an improved diffusion model for RGBX, which also estimates lighting, as well as the first diffusion XRGB model capable of synthesizing realistic images from (full or partial) intrinsic channels. Our XRGB model explores a middle ground between traditional rendering and generative models: we can specify only certain appearance properties that should be followed, and give freedom to the model to hallucinate a plausible version of the rest. This flexibility makes it possible to use a mix of heterogeneous training datasets, which differ in the available channels. We use multiple existing datasets and extend them with our own synthetic and real data, resulting in a model capable of extracting scene properties better than previous work and of generating highly realistic images of interior scenes.
References in corpus (4)
Cited by in corpus (7)
- IntrinsicEdit: Precise generative image manipulation in intrinsic space
- LightLab: Controlling Light Sources in Images with Diffusion Models
- Generative Detail Enhancement for Physically Based Materials
- Uncertainty for SVBRDF Acquisition using Frequency Analysis
- MaterialPicker: Multi-Modal DiT-Based Material Generation
- Aerial Path Planning for Urban Geometry and Texture Co-Capture
- Detail Enhanced Gaussian Splatting for Large-Scale Volumetric Capture