activity
20242026
collaborators

22 papers

cs.CV2026

AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT

Guandong Li, Mengxia Ye

We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-image tokens within a single a…

cs.CV2026

PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning

Guandong Li, Mengxia Ye

Image editing instructions are heterogeneous: a color swap, an object insertion, and a physical-action edit all demand different spatial coverage and different reasoning depth, yet…

cs.CV2026

Edit Fidelity Field: Semantics-Aware Region Isolation for Training-Free Scene Text Editing

Guandong Li, Mengxia Ye

Scene text editing (STE) has achieved remarkable progress in accurately rendering target text through diffusion-based methods. However, we identify a critical yet overlooked proble…

cs.CV2026

LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference

Guandong Li

Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Transformer networks. We observ…

cs.CV2026

AdaEdit: Adaptive Temporal and Channel Modulation for Flow-Based Image Editing

Guandong Li, Zhaobin Chu

Inversion-based image editing in flow matching models has emerged as a powerful paradigm for training-free, text-guided image manipulation. A central challenge in this paradigm is…

cs.CV2026

Edit Spillover as a Probe: Do Image Editing Models Implicitly Understand World Relations?

Guandong Li, Zhaobin Chu

Instruction-following image editing models are expected to modify only the specified region while keeping the rest of the image unchanged. However, in practice, we observe a pervas…