4D Facial Expression Diffusion Model
arXiv:2303.16611 · doi:10.1145/3653455
Abstract
Facial expression generation is one of the most challenging and long-sought aspects of character animation, with many interesting applications. The challenging task, traditionally having relied heavily on digital craftspersons, remains yet to be explored. In this paper, we introduce a generative framework for generating 3D facial expression sequences (i.e. 4D faces) that can be conditioned on different inputs to animate an arbitrary 3D face mesh. It is composed of two tasks: (1) Learning the generative model that is trained over a set of 3D landmark sequences, and (2) Generating 3D mesh sequences of an input facial mesh driven by the generated landmark sequences. The generative model is based on a Denoising Diffusion Probabilistic Model (DDPM), which has achieved remarkable success in generative tasks of other domains. While it can be trained unconditionally, its reverse process can still be conditioned by various condition signals. This allows us to efficiently develop several downstream tasks involving various conditional generation, by using expression labels, text, partial sequences, or simply a facial geometry. To obtain the full mesh deformation, we then develop a landmark-guided encoder-decoder to apply the geometrical deformation embedded in landmarks on a given facial mesh. Experiments show that our model has learned to generate realistic, quality expressions solely from the dataset of relatively small size, improving over the state-of-the-art methods. Videos and qualitative comparisons with other methods can be found at \url{https://github.com/ZOUKaifeng/4DFM}.
References in corpus (23)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- NICE: Non-linear Independent Components Estimation
- Score-Based Generative Modeling through Stochastic Differential Equations
- Cascaded Diffusion Models for High Fidelity Image Generation
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Action2Motion: Conditioned Generation of 3D Human Motions
- Imagen Video: High Definition Video Generation with Diffusion Models
- Diffusion-LM Improves Controllable Text Generation
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Flexible Diffusion Modeling of Long Videos
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
- Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
- Structured Denoising Diffusion Models in Discrete State-Spaces
- Diffusion-based Time Series Imputation and Forecasting with Structured State Space Models
- Image Super-Resolution via Iterative Refinement
- DiffuseVAE: Efficient, Controllable and High-Fidelity Generation from Low-Dimensional Latents
- Label-Efficient Semantic Segmentation with Diffusion Models
- Facial Expression Video Generation Based-On Spatio-temporal Convolutional GAN: FEV-GAN
- A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion
- EDICT: Exact Diffusion Inversion via Coupled Transformations