3 papers
cs.GR2026
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
Aryan Mikaeili, Or Patashnik, Andrea Tagliasacchi +2
Positional encodings are essential to transformer-based generative models, yet their behavior in multimodal and attention-sharing settings is not fully understood. In this work, we…
cs.CV2025
Griffin: Generative Reference and Layout Guided Image Composition
Aryan Mikaeili, Amirhossein Alimohammadi, Negar Hassanpour +2
Text-to-image models have achieved a level of realism that enables the generation of highly convincing images. However, text-based control can be a limiting factor when more explic…
cs.CV2025
Cora: Correspondence-aware image editing using few step diffusion
Amirhossein Alimohammadi, Aryan Mikaeili, Sauradip Nag +3
Image editing is an important task in computer graphics, vision, and VFX, with recent diffusion-based methods achieving fast and high-quality results. However, edits requiring sign…