Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
arXiv:2306.12760 · doi:10.1109/ICCVW60793.2023.00316
Abstract
Editing a local region or a specific object in a 3D scene represented by a NeRF or consistently blending a new realistic object into the scene is challenging, mainly due to the implicit nature of the scene representation. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.
16 pages, 14 figures. Project page: https://www.vision.huji.ac.il/blended-nerf/
References in corpus (9)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis
- Language-driven Semantic Segmentation
- MetaSDF: Meta-learning Signed Distance Functions
- CIPS-3D: A 3D-Aware Generator of GANs Based on Conditionally-Independent Pixel Synthesis
- UPST-NeRF: Universal Photorealistic Style Transfer of Neural Radiance Fields for 3D Scene