3 papers
eess.AS2026
Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Yunjia Li, Menglin Wu, Junyu Dai +13
Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…
cs.CV2025
Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
Haiyang Liu, Xiaolin Hong, Xuancheng Yang +5
We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We addres…
cs.CV2025
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
Hongwei Yi, Tian Ye, Shitong Shao +10
We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse c…