1 paper
Chunyu Li, Jiaye Li, Ruiqiao Mei +4
Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio…