HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances
arXiv:2403.01693 · doi:10.1109/CVPR52733.2024.00239
Abstract
Text-to-image generative models can generate high-quality humans, but realism is lost when generating hands. Common artifacts include irregular hand poses, shapes, incorrect numbers of fingers, and physically implausible finger orientations. To generate images with realistic hands, we propose a novel diffusion-based architecture called HanDiffuser that achieves realism by injecting hand embeddings in the generative process. HanDiffuser consists of two components: a Text-to-Hand-Params diffusion model to generate SMPL-Body and MANO-Hand parameters from input text prompts, and a Text-Guided Hand-Params-to-Image diffusion model to synthesize images by conditioning on the prompts and hand parameters generated by the previous component. We incorporate multiple aspects of hand representation, including 3D shapes and joint-level finger positions, orientations and articulations, for robust learning and reliable performance during inference. We conduct extensive quantitative and qualitative experiments and perform user studies to demonstrate the efficacy of our method in generating images with high-quality hands.
Revisions: 1. Added a link to project page in the abstract, 2. Updated references and related work, 3. Fixed some grammatical errors
References in corpus (10)
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Embodied Hands: Modeling and Capturing Hands and Bodies Together
- MediaPipe Hands: On-device Real-time Hand Tracking
- Cascaded Diffusion Models for High Fidelity Image Generation
- The KIT Motion-Language Dataset
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- Detecting Hands and Recognizing Physical Contact in the Wild
- GRIP: Generating Interaction Poses Using Spatial Cues and Latent Consistency
- HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting