5 papers
Semantic Foam: Unifying Spatial and Semantic Scene Decomposition
Amr Sharafeldin, Shrisudhan Govindarajan, Thomas Walker +4
Modern scene reconstruction methods, such as 3D Gaussian Splatting, deliver photo-realistic novel view synthesis at real-time speeds, yet their adoption in interactive graphics app…
Sound Sparks Motion: Audio and Text Tuning for Video Editing
AmirHossein Naghi Razlighi, Aryan Mikaeili, Ali Mahdavi-Amiri +2
Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions…
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
Aryan Mikaeili, Or Patashnik, Andrea Tagliasacchi +2
Positional encodings are essential to transformer-based generative models, yet their behavior in multimodal and attention-sharing settings is not fully understood. In this work, we…
Griffin: Generative Reference and Layout Guided Image Composition
Aryan Mikaeili, Amirhossein Alimohammadi, Negar Hassanpour +2
Text-to-image models have achieved a level of realism that enables the generation of highly convincing images. However, text-based control can be a limiting factor when more explic…
Cora: Correspondence-aware image editing using few step diffusion
Amirhossein Alimohammadi, Aryan Mikaeili, Sauradip Nag +3
Image editing is an important task in computer graphics, vision, and VFX, with recent diffusion-based methods achieving fast and high-quality results. However, edits requiring sign…