ESCT3D: Efficient and Selectively Controllable Text-Driven 3D Content Generation with Gaussian Splatting
arXiv:2504.10316 · doi:10.1145/3728305
Abstract
In recent years, significant advancements have been made in text-driven 3D content generation. However, several challenges remain. In practical applications, users often provide extremely simple text inputs while expecting high-quality 3D content. Generating optimal results from such minimal text is a difficult task due to the strong dependency of text-to-3D models on the quality of input prompts. Moreover, the generation process exhibits high variability, making it difficult to control. Consequently, multiple iterations are typically required to produce content that meets user expectations, reducing generation efficiency. To address this issue, we propose GPT-4V for self-optimization, which significantly enhances the efficiency of generating satisfactory content in a single attempt. Furthermore, the controllability of text-to-3D generation methods has not been fully explored. Our approach enables users to not only provide textual descriptions but also specify additional conditions, such as style, edges, scribbles, poses, or combinations of multiple conditions, allowing for more precise control over the generated 3D content. Additionally, during training, we effectively integrate multi-view information, including multi-view depth, masks, features, and images, to address the common Janus problem in 3D content generation. Extensive experiments demonstrate that our method achieves robust generalization, facilitating the efficient and controllable generation of high-quality 3D content.
References in corpus (37)
- Denoising Diffusion Probabilistic Models
- Instant Neural Graphics Primitives with a Multiresolution Hash Encoding
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- DINOv2: Learning Robust Visual Features without Supervision
- Classifier-Free Diffusion Guidance
- Extended Bayesian Information Criteria for Gaussian Graphical Models
- DreamFusion: Text-to-3D using 2D Diffusion
- Blended Latent Diffusion
- Implicit Neural Representations with Periodic Activation Functions
- CLIP-Mesh: Generating textured meshes from text using pretrained image-text models
- LION: Latent Point Diffusion Models for 3D Shape Generation
- GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images
- Point-E: A System for Generating 3D Point Clouds from Complex Prompts
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
- Shap-E: Generating Conditional 3D Implicit Functions
- PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
- Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape Synthesis
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
- Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors
- Neural Point Catacaustics for Novel-View Synthesis of Reflections
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation
- MVDream: Multi-view Diffusion for 3D Generation
- Optimizing Prompts for Text-to-Image Generation
- Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
- 3DGen: Triplane Latent Diffusion for Textured Mesh Generation
- MeshDiffusion: Score-based Generative 3D Mesh Modeling
- DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior
- DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
- SweetDreamer: Aligning Geometric Priors in 2D Diffusion for Consistent Text-to-3D
- LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes
- ED-NeRF: Efficient Text-Guided Editing of 3D Scene with Latent Space NeRF
- ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
- Controllable Text-to-3D Generation via Surface-Aligned Gaussian Splatting
- RoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture