5 papers
StreamDiT: Real-Time Streaming Text-to-Video Generation
Akio Kodaira, Tingbo Hou, Ji Hou +4
Recently, great progress has been achieved in text-to-video (T2V) generation by scaling transformer-based diffusion models to billions of parameters, which can generate high-qualit…
StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
Tianrui Feng, Zhi Li, Shuo Yang +11
Generative models are reshaping the live-streaming industry by redefining how content is created, styled, and delivered. Previous image-based streaming diffusion models have powere…
StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation
Akio Kodaira, Chenfeng Xu, Toshiki Hazama +8
We introduce StreamDiffusion, a real-time diffusion pipeline designed for interactive image generation. Existing diffusion models are adept at creating images from text or image pr…
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
Feng Liang, Akio Kodaira, Chenfeng Xu +3
This paper introduces StreamV2V, a diffusion model that achieves real-time streaming video-to-video (V2V) translation with user prompts. Unlike prior V2V methods using batches to p…
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
Yiheng Li, Heyang Jiang, Akio Kodaira +3
In this paper, we point out that suboptimal noise-data mapping leads to slow training of diffusion models. During diffusion training, current methods diffuse each image across the…