28 citations · 92 across the 13 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 11 cited
UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance
Wei Li, Xue Xu, Xinyan Xiao +8
Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusio…
cs.CV2022★ 2 cited
UNIMO-2: End-to-End Unified Vision-Language Grounded Learning
Wei Li, Can Gao, Guocheng Niu +5
Vision-Language Pre-training (VLP) has achieved impressive performance on various cross-modal downstream tasks. However, most existing methods can only learn from aligned image-cap…