10 papers
Detail++: Training-Free Detail Enhancer for T2I Diffusion Models
Lifeng Chen, Jiner Wang, Zihao Pan +3
Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, parti…
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
Zhengrong Yue, Taihang Hu, Mengting Chen +8
Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designe…
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
Tao Liu, Hao Yan, Mengting Chen +8
Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two repres…
Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation
Daiqiang Li, Zihao Pan, Zeyu Zhang +8
In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
Zhangquan Chen, Manyuan Zhang, Xinlei Yu +8
Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limit…
X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
Jian Ma, Xujie Zhu, Zihao Pan +4
Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models…