83 citations · 179 across the 13 of their papers we have counts for
21 papers
Can Vision Transformers Learn without Natural Images?
Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto +2
Can we complete pre-training of Vision Transformers (ViT) without natural images and human-annotated labels? Although a pre-trained ViT seems to heavily rely on a large-scale datas…
Describing and Localizing Multiple Changes with Transformers
Yue Qiu, Shintaro Yamamoto, Kodai Nakashima +4
Change captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change ca…
Pre-training without Natural Images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto +5
Is it possible to use convolutional neural networks pre-trained without any natural images to assist natural image understanding? The paper proposes a novel concept, Formula-driven…
Initialization Using Perlin Noise for Training Networks with a Limited Amount of Data
Nakamasa Inoue, Eisuke Yamagata, Hirokatsu Kataoka
We propose a novel network initialization method using Perlin noise for training image classification networks with a limited amount of data. Our main idea is to initialize the net…
Alleviating Over-segmentation Errors by Detecting Action Boundaries
Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki +1
We propose an effective framework for the temporal action segmentation task, namely an Action Segment Refinement Framework (ASRF). Our model architecture consists of a long-term fe…
Retrieving and Highlighting Action with Spatiotemporal Reference
Seito Kasai, Yuchi Ishikawa, Masaki Hayashi +3
In this paper, we present a framework that jointly retrieves and spatiotemporally highlights actions in videos by enhancing current deep cross-modal retrieval methods. Our work tak…