collaborators

5 papers

cs.CV2026

PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation

Yanjie Pan, Qingdong He, Zhengkai Jiang +9

Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle…

cs.CV2026

Towards Generalized Multi-Image Editing for Unified Multimodal Models

Pengcheng Xu, Peng Tang, Donghao Luo +7

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when refer…

cs.CV2025

DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation

Qingdong He, Jinlong Peng, Pengcheng Xu +8

To enhance the controllability of text-to-image diffusion models, current ControlNet-like models have explored various control signals to dictate image attributes. However, existin…

cs.CV2025

VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models

Hu Xiaobin, Liang Yujie, Luo Donghao +5

While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three k…

cs.CV2025

Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

Pengcheng Xu, Boyuan Jiang, Xiaobin Hu +7

Leveraging the large generative prior of the flow transformer for tuning-free image editing requires authentic inversion to project the image into the model's domain and a flexible…