11 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 1 cited
CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
Guohao Li, Hu Yang, Feng He +4
Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner,…
cs.CV2022★ 11 cited
UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance
Wei Li, Xue Xu, Xinyan Xiao +8
Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusio…