1 paper
Yubo Huang, Sirui Zhao, Xinchen Yao +5
Audio-driven avatar generation requires realistic lip-sync, expressive motion, and real-time streaming. Recent work achieves the latter via self-forcing with Distribution Matching…