9 papers
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
Yoav Baron, Sara Dorfman, Roni Paiss +2
Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarkable multimodal reasoning capabilities. Wh…
Versatile Editing of Video Content, Actions, and Dynamics without Training
Vladimir Kulikov, Roni Paiss, Andrey Voynov +3
Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of ot…
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
Ronen Kamenetsky, Sara Dorfman, Daniel Garibi +3
Large-scale text-to-image diffusion models have become the backbone of modern image editing, yet text prompts alone do not offer adequate control over the editing process. Two prop…
What Are You Doing? A Closer Look at Controllable Human Video Generation
Emanuele Bugliarello, Anurag Arnab, Roni Paiss +2
High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset…
TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space
Daniel Garibi, Shahar Yadin, Roni Paiss +6
We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements a…
Imagen 3
Imagen-Team-Google, :, Jason Baldridge +257
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…