3 papers
cs.CV2024
Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
Shiyin Dong, Mingrui Zhu, Kun Cheng +2
The remarkable prowess of diffusion models in image generation has spurred efforts to extend their application beyond generative tasks. However, a persistent challenge exists in la…
cs.CV2023
CatVersion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization
Ruoyu Zhao, Mingrui Zhu, Shiyin Dong +2
We propose CatVersion, an inversion-based method that learns the personalized concept through a handful of examples. Subsequently, users can utilize text prompts to generate images…
cs.CV2023
Adapt and Align to Improve Zero-Shot Sketch-Based Image Retrieval
Shiyin Dong, Mingrui Zhu, Nannan Wang +1
Zero-shot sketch-based image retrieval (ZS-SBIR) is challenging due to the cross-domain nature of sketches and photos, as well as the semantic gap between seen and unseen image dis…