2 papers
cs.CV2024
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
Zilyu Ye, Jinxiu Liu, Ruotian Peng +9
Recent image generation models excel at creating high-quality images from brief captions. However, they fail to maintain consistency of multiple instances across images when encoun…
cs.CV2023
AdPE: Adversarial Positional Embeddings for Pretraining Vision Transformers via MAE+
Xiao Wang, Ying Wang, Ziwei Xuan +1
Unsupervised learning of vision transformers seeks to pretrain an encoder via pretext tasks without labels. Among them is the Masked Image Modeling (MIM) aligned with pretraining o…