1 paper · 1 filter
William Berman, Alexander Peysakhovich
We train a model to generate images from multimodal prompts of interleaved text and images such as "a <picture of a man> man and his <picture of a dog> dog in an <picture of a cart…