7 citations · 14 across the 2 of their papers we have counts for
7 papers
Perspectives and Prospects on Transformer Architecture for Cross-Modal Tasks with Language and Vision
Andrew Shin, Masato Ishii, Takuya Narihira
Transformer architectures have brought about fundamental changes to computational linguistic field, which had been dominated by recurrent neural networks for many years. Its succes…
Neural Network Libraries: A Deep Learning Framework Designed from Engineers' Perspectives
Takuya Narihira, Javier Alonsogarcia, Fabien Cardinaux +14
While there exist a plethora of deep learning tools and frameworks, the fast-growing complexity of the field brings new demands and challenges, such as more flexible network design…
Reference-Based Video Colorization with Spatiotemporal Correspondence
Naofumi Akimoto, Akio Hayakawa, Andrew Shin +1
We propose a novel reference-based video colorization framework with spatiotemporal correspondence. Reference-based methods colorize grayscale frames referencing a user input color…
Customized Image Narrative Generation via Interactive Visual Question Generation and Answering
Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada
Image description task has been invariably examined in a static manner with qualitative presumptions held to be universally applicable, regardless of the scope or target of the des…
Melody Generation for Pop Music via Word Representation of Musical Properties
Andrew Shin, Leopold Crestel, Hiroharu Kato +6
Automatic melody generation for pop music has been a long-time aspiration for both AI researchers and musicians. However, learning to generate euphonious melody has turned out to b…
Beyond Caption To Narrative: Video Captioning With Multiple Sentences
Andrew Shin, Katsunori Ohnishi, Tatsuya Harada
Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input o…