CLIP-Mesh: Generating textured meshes from text using pretrained image-text models
arXiv:2203.13333 · doi:10.1145/3550469.3555392
Abstract
We present a technique for zero-shot generation of a 3D model using only a target text prompt. Without any 3D supervision our method deforms the control shape of a limit subdivided surface along with its texture map and normal map to obtain a 3D asset that corresponds to the input text prompt and can be easily deployed into games or modeling applications. We rely only on a pre-trained CLIP model that compares the input text prompt with differentiably rendered images of our 3D model. While previous works have focused on stylization or required training of generative models we perform optimization on mesh parameters directly to generate shape, texture or both. To constrain the optimization to produce plausible meshes and textures we introduce a number of techniques using image augmentations and the use of a pretrained prior that generates CLIP image embeddings given a text embedding.
8 pages, 8 figures, Accepted at SIGGRAPH ASIA 2022, Project Page at https://www.nasir.lol/clipmesh
References in corpus (2)
Cited by in corpus (9)
- Text-Guided Texturing by Synchronized Multi-View Diffusion
- FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications
- SHAPE-IT: Exploring Text-to-Shape-Display for Generative Shape-Changing Behaviors with LLMs
- Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
- Localized Gaussian Splatting Editing with Contextual Awareness
- GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts
- Controllable 3D Object Generation with Single Image Prompt
- ESCT3D: Efficient and Selectively Controllable Text-Driven 3D Content Generation with Gaussian Splatting
- Advances in Neural 3D Mesh Texturing: A Survey