collaborators

10 papers

cs.CV2026

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

Lifeng Chen, Jiner Wang, Zihao Pan +3

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, parti…

cs.CV2026

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

Zhengrong Yue, Taihang Hu, Mengting Chen +8

Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designe…

cs.CV2026

Continuous-Time Distribution Matching for Few-Step Diffusion Distillation

Tao Liu, Hao Yan, Mengting Chen +8

Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two repres…

cs.CV2026

Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation

Daiqiang Li, Zihao Pan, Zeyu Zhang +8

In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…

cs.CV2026

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views

Zhangquan Chen, Manyuan Zhang, Xinlei Yu +8

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limit…

cs.CV2025

X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning

Jian Ma, Xujie Zhu, Zihao Pan +4

Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models…