8 papers
SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting
Hongzhe Liao, Chuhua Xian, Hongmin Cai +2
3D Gaussian Splatting (3DGS) provides an efficient method for high-quality scene reconstruction using anisotropic Gaussians. Recently, 3DGS-based methods have significantly improve…
TDMM-LM: Bridging Facial Understanding and Animation via Language Models
Luchuan Song, Pinxin Liu, Haiyang Liu +7
Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage f…
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
Bohong Chen, Haiyang Liu
Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. Th…
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
Pinxin Liu, Haiyang Liu, Luchuan Song +2
When humans speak, gestures help convey communicative intentions, such as adding emphasis or describing concepts. However, current co-speech gesture generation methods rely solely…
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
Xiaochuan Li, Guoguang Du, Runze Zhang +11
Scaling laws have validated the success and promise of large-data-trained models in creative generation across text, image, and video domains. However, this paradigm faces data sca…
Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
Haiyang Liu, Xiaolin Hong, Xuancheng Yang +5
We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We addres…