22 papers
AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT
Guandong Li, Mengxia Ye
We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-image tokens within a single a…
PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning
Guandong Li, Mengxia Ye
Image editing instructions are heterogeneous: a color swap, an object insertion, and a physical-action edit all demand different spatial coverage and different reasoning depth, yet…
Edit Fidelity Field: Semantics-Aware Region Isolation for Training-Free Scene Text Editing
Guandong Li, Mengxia Ye
Scene text editing (STE) has achieved remarkable progress in accurately rendering target text through diffusion-based methods. However, we identify a critical yet overlooked proble…
LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference
Guandong Li
Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Transformer networks. We observ…
AdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing
Guandong Li, Zhaobin Chu
Inversion-based image editing in flow matching models has emerged as a powerful paradigm for training-free, text-guided image manipulation. A central challenge in this paradigm is…
Edit Spillover as a Probe: Do Image Editing Models Implicitly Understand World Relations?
Guandong Li, Zhaobin Chu
Instruction-following image editing models are expected to modify only the specified region while keeping the rest of the image unchanged. However, in practice, we observe a pervas…