1 citations · 3 across the 37 of their papers we have counts for
11 papers · 1 filter
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
Aryan Mikaeili, Or Patashnik, Andrea Tagliasacchi +2
Positional encodings are essential to transformer-based generative models, yet their behavior in multimodal and attention-sharing settings is not fully understood. In this work, we…
LooseRoPE: Content-aware Attention Manipulation for Semantic Harmonization
Etai Sella, Yoav Baron, Hadar Averbuch-Elor +2
Recent diffusion-based image editing methods commonly rely on text or high-level instructions to guide the generation process, offering intuitive but coarse control. In contrast, w…
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
Ronen Kamenetsky, Sara Dorfman, Daniel Garibi +3
Large-scale text-to-image diffusion models have become the backbone of modern image editing, yet text prompts alone do not offer adequate control over the editing process. Two prop…
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
Rameen Abdal, Or Patashnik, Ekaterina Deyneka +5
Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-sp…
EditP23: 3D Editing via Propagation of Image Prompts to Multi-View
Roi Bar-On, Dana Cohen-Bar, Daniel Cohen-Or
We present EditP23, a method for mask-free 3D editing that propagates 2D image edits to multi-view representations in a 3D-consistent manner. In contrast to traditional approaches…
Navigating with Annealing Guidance Scale in Diffusion Space
Shai Yehezkel, Omer Dahary, Andrey Voynov +1
Denoising diffusion models excel at generating high-quality images conditioned on text prompts, yet their effectiveness heavily relies on careful guidance during the sampling proce…