5 papers
Diffusion Image Editing via Asynchronous Token Decoding
Yang Shi, Liangsi Lu, Minzhe Guo +4
Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condit…
Semantic Granularity Navigation in Image Editing
Liangsi Lu, Minzhe Guo, Xuhang Chen +1
Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic editability and structural fidel…
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
Yang Shi, Yifeng Xie, Minzhe Guo +6
Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models truly understand the content they p…
ChordEdit: One-Step Low-Energy Transport for Image Editing
Liangsi Lu, Xuhang Chen, Minzhe Guo +3
The advent of one-step text-to-image (T2I) models offers unprecedented synthesis speed. However, their application to text-guided image editing remains severely hampered, as forcin…
PC-UNet: An Enforcing Poisson Statistics U-Net for Positron Emission Tomography Denoising
Yang Shi, Jingchao Wang, Liangsi Lu +9
Positron Emission Tomography (PET) is crucial in medicine, but its clinical use is limited due to high signal-to-noise ratio doses increasing radiation exposure. Lowering doses inc…