7 papers
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
Hila Manor, Rinon Gal, Haggai Maron +2
Visual analogy learning enables image editing via demonstration rather than textual description, allowing users to specify complex transformations difficult to articulate in words.…
ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation
Rotem Shalev-Arkushin, Rinon Gal, Amit H. Bermano +1
Diffusion models enable high-quality and diverse visual content synthesis. However, they struggle to generate rare or unseen concepts. To address this challenge, we explore the usa…
Policy Optimized Text-to-Image Pipeline Design
Uri Gadot, Rinon Gal, Yftah Ziser +2
Text-to-image generation has evolved beyond single monolithic models to complex multi-component pipelines. These combine fine-tuned generators, adapters, upscaling blocks and even…
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
Yuval Atzmon, Rinon Gal, Yoad Tewel +2
Text-to-video diffusion models have shown remarkable progress in generating coherent video clips from textual descriptions. However, the interplay between motion, structure, and id…
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
Michael Toker, Ido Galil, Hadas Orgad +4
Text-to-image (T2I) diffusion models rely on encoded prompts to guide the image generation process. Typically, these prompts are extended to a fixed length by adding padding tokens…
IP-Composer: Semantic Composition of Visual Concepts
Sara Dorfman, Dana Cohen-Bar, Rinon Gal +1
Content creators often draw inspiration from multiple visual sources, combining distinct elements to craft new compositions. Modern computational approaches now aim to emulate this…