7 papers
Diffusion Image Editing via Asynchronous Token Decoding
Yang Shi, Liangsi Lu, Minzhe Guo +4
Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condit…
Semantic Granularity Navigation in Image Editing
Liangsi Lu, Minzhe Guo, Xuhang Chen +1
Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic editability and structural fidel…
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
Yang Shi, Yifeng Xie, Minzhe Guo +6
Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models truly understand the content they p…
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
Alizer Wong, Zixin Zeng, Yi Tan +6
Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial mathematic…
ChordEdit: One-Step Low-Energy Transport for Image Editing
Liangsi Lu, Xuhang Chen, Minzhe Guo +3
The advent of one-step text-to-image (T2I) models offers unprecedented synthesis speed. However, their application to text-guided image editing remains severely hampered, as forcin…
Riemannian Liquid Spatio-Temporal Graph Network
Liangsi Lu, Jingchao Wang, Zhaorong Dai +2
Liquid Time-Constant networks (LTCs), a type of continuous-time graph neural network, excel at modeling irregularly-sampled dynamics but are fundamentally confined to Euclidean spa…