From the 1 of 12 linked papers with an AI index.
12 papers
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Yixuan Lai, Tianjia Shao, Kun Zhou +3
GroundShot is a training-free, model-agnostic framework that improves visual consistency in multi-shot video generation by maintaining an online entity-level visual memory and sche…
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
Hui Li, Fu-Yun Wang, Haoyuan Xia +4
This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one tra…
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation
Yuxuan Yao, Yuxuan Chen, Hui Li +6
Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visua…
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
Weijia Dou, Hui Li, Jiahao Cui +3
Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclustered tokens. This organizat…
Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds
Liuzhuozheng Li, Zhiyuan Zhan, Shuhong Liu +5
Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Especially in deterministic method…
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
Chunyu Li, Jiaye Li, Ruiqiao Mei +4
Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio…