activity
20242026
collaborators

9 papers

cs.CV2026

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

Yoav Baron, Sara Dorfman, Roni Paiss +2

Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarkable multimodal reasoning capabilities. Wh…

cs.CV2026

Versatile Editing of Video Content, Actions, and Dynamics without Training

Vladimir Kulikov, Roni Paiss, Andrey Voynov +3

Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of ot…

cs.GR2025

SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder

Ronen Kamenetsky, Sara Dorfman, Daniel Garibi +3

Large-scale text-to-image diffusion models have become the backbone of modern image editing, yet text prompts alone do not offer adequate control over the editing process. Two prop…

cs.CV2025

What Are You Doing? A Closer Look at Controllable Human Video Generation

Emanuele Bugliarello, Anurag Arnab, Roni Paiss +2

High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset…

cs.CV2025

TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space

Daniel Garibi, Shahar Yadin, Roni Paiss +6

We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements a…

cs.CV2024

Imagen 3

Imagen-Team-Google, :, Jason Baldridge +257

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…