4 papers
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
Yuanbing Zhu, Bingke Zhu, Yingying Chen +3
Pretrained vision-language models (VLMs), \eg CLIP, are increasingly used to bridge the gap between open- and close-vocabulary recognition in open-vocabulary image segmentation. As…
AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
Yunfang Niu, Dong Yi, Lingxiang Wu +2
Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters and keypoint extractors, lacking a…
Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
Pengxiang Cai, Zhiwei Liu, Guibo Zhu +2
Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fai…
PFDM: Parser-Free Virtual Try-on via Diffusion Model
Yunfang Niu, Dong Yi, Lingxiang Wu +3
Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve h…