17 papers · 1 filter
EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing
Yuqian Zhou, Zhenghong Zhou, Zongze Wu +5
Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video…
HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing
Haoran You, Yotam Nitzan, Lingzhi Zhang +8
Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a major share of traffic in Photoshop and…
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen +8
Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for c…
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
Haoyu Chen, Qing Liu, Yuqian Zhou +7
Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffe…
Rethinking Global Text Conditioning in Diffusion Transformers
Nikita Starodubcev, Daniil Pakhomov, Zongze Wu +6
Diffusion transformers typically incorporate textual information via attention layers and a modulation mechanism using a pooled text embedding. Nevertheless, recent approaches disc…
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Shilong Zhang, He Zhang, Zhifei Zhang +11
Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To uni…