12 papers
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
Dongting Hu, Aarush Gupta, Magzhan Gabidolla +12
Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and mem…
Prompt2Effect: Training-Free Image-to-Video Model Specialization via LoRA Generation
Xiaomeng Yang, Yanyu Li, Gordon Guocheng Qian +7
While personalizing Image-to-Video (I2V) diffusion models with specific visual effects is increasingly demanded for high-end generation, current practice requires training a separa…
GeoStream: Toward Precise Camera Controlled Streaming Video Generation
Yizhou Zhao, Yifan Wang, Xiaoyuan Wang +11
Accurate interactive camera control is essential for video-based world models, but most existing approaches learn camera motion implicitly, leading to inaccurate control under out-…
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
Dogyun Park, Yanyu Li, Sergey Tulyakov +1
Scaling video diffusion transformers is fundamentally bottlenecked by two compounding costs: the expensive quadratic complexity of attention per step, and the iterative sampling st…
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
Yifan Wang, Yanyu Li, Gordon Guocheng Qian +3
Video diffusion alignment has been heavily relied on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring addition…
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
Lin Zhao, Yushu Wu, Aleksei Lebedev +11
Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this w…