2 papers
cs.CV2021
Can Vision Transformers Learn without Natural Images?
Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto +2
Can we complete pre-training of Vision Transformers (ViT) without natural images and human-annotated labels? Although a pre-trained ViT seems to heavily rely on a large-scale datas…
cs.CV2021
Describing and Localizing Multiple Changes with Transformers
Yue Qiu, Shintaro Yamamoto, Kodai Nakashima +4
Change captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change ca…