4 papers
Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Yunjia Li, Menglin Wu, Junyu Dai +13
Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are tra…
SDFlow: Similarity-Driven Flow Matching for Time Series Generation
Wei Li, Shibo Feng, Pengcheng Wu +3
Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamenta…
Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
Haiyang Liu, Xiaolin Hong, Xuancheng Yang +5
We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We addres…
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
Hongwei Yi, Tian Ye, Shitong Shao +10
We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse c…