2 papers
cs.CV2024
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
Ming Tao, Bing-Kun Bao, Yaowei Wang +1
Large pretrained diffusion models have demonstrated impressive generation capabilities and have been adapted to various downstream tasks. However, unlike Large Language Models (LLM…
cs.CV2024
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
Ming Tao, Bing-Kun Bao, Hao Tang +2
Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trai…