5 papers
ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation
Rotem Shalev-Arkushin, Rinon Gal, Amit H. Bermano +1
Diffusion models enable high-quality and diverse visual content synthesis. However, they struggle to generate rare or unseen concepts. To address this challenge, we explore the usa…
V-LASIK: Consistent Glasses-Removal from Videos Using Synthetic Data
Rotem Shalev-Arkushin, Aharon Azulay, Tavi Halperin +3
Diffusion-based generative models have recently shown remarkable image and video editing capabilities. However, local video editing, particularly removal of small attributes like g…
Ham2Pose: Animating Sign Language Notation into Pose Sequences
Rotem Shalev-Arkushin, Amit Moryossef, Ohad Fried
Translating spoken languages into Sign languages is necessary for open communication between the hearing and hearing-impaired communities. To achieve this goal, we propose the firs…
Stable Flow: Vital Layers for Training-Free Image Editing
Omri Avrahami, Or Patashnik, Ohad Fried +4
Diffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional UNet architecture with the Diffusion Transformer (DiT),…
Advancing Fine-Grained Classification by Structure and Subject Preserving Augmentation
Eyal Michaeli, Ohad Fried
Fine-grained visual classification (FGVC) involves classifying closely related sub-classes. This task is difficult due to the subtle differences between classes and the high intra-…