4 papers · 1 filter
Guidelines for External Disturbance Factors in the Use of OCR in Real-World Environments
Kenji Iwata, Eiki Ishidera, Toshifumi Yamaai +6
The performance of OCR has improved with the evolution of AI technology. As OCR continues to broaden its range of applications, the increased likelihood of interference introduced…
Can Vision Transformers Learn without Natural Images?
Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto +2
Can we complete pre-training of Vision Transformers (ViT) without natural images and human-annotated labels? Although a pre-trained ViT seems to heavily rely on a large-scale datas…
Describing and Localizing Multiple Changes with Transformers
Yue Qiu, Shintaro Yamamoto, Kodai Nakashima +4
Change captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change ca…
Dominant Codewords Selection with Topic Model for Action Recognition
Hirokatsu Kataoka, Masaki Hayashi, Kenji Iwata +3
In this paper, we propose a framework for recognizing human activities that uses only in-topic dominant codewords and a mixture of intertopic vectors. Latent Dirichlet allocation (…