8 papers
Diffusion Image Editing via Asynchronous Token Decoding
Yang Shi, Liangsi Lu, Minzhe Guo +4
Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condit…
TransSplat: Unbalanced Semantic Transport for Language-Driven 3DGS Editing
Yanhui Chen, Jiahong Li, Jingchao Wang +3
Language-driven 3D Gaussian Splatting (3DGS) editing provides a more convenient approach for modifying complex scenes in VR/AR. Standard pipelines typically adopt a two-stage strat…
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
Yanhui Chen, Baoyao Yang, Siqi Liu +1
SAM3 advances open-vocabulary semantic segmentation by introducing a prompt-driven mask generation paradigm. However, in multi-class open-vocabulary scenarios, masks generated inde…
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
Yang Shi, Yifeng Xie, Minzhe Guo +6
Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models truly understand the content they p…
ChordEdit: One-Step Low-Energy Transport for Image Editing
Liangsi Lu, Xuhang Chen, Minzhe Guo +3
The advent of one-step text-to-image (T2I) models offers unprecedented synthesis speed. However, their application to text-guided image editing remains severely hampered, as forcin…
Riemannian Liquid Spatio-Temporal Graph Network
Liangsi Lu, Jingchao Wang, Zhaorong Dai +2
Liquid Time-Constant networks (LTCs), a type of continuous-time graph neural network, excel at modeling irregularly-sampled dynamics but are fundamentally confined to Euclidean spa…