Visual Conceptual Blending with Large-scale Language and Vision Models
arXiv:2106.14127
Abstract
We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description of the blend of the two using a language model. We then generate a visual depiction of the blend using a text-based image generation model. Quantitative and qualitative evaluations demonstrate the superiority of language models over classical methods for conceptual blending, and of recent large-scale image generation models over prior models for the visual depiction.
References in corpus (5)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Zero-Shot Text-to-Image Generation
- DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis
- Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile Generation