activity
20202024
most citedAudio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

3 citations · 9 across the 15 of their papers we have counts for

collaborators

14 papers

cs.CL2022

On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment Analysis

Atsushi Ando, Ryo Masumura, Akihiko Takashima +6

This paper investigates the effectiveness and implementation of modality-specific large-scale pre-trained encoders for multimodal sentiment analysis~(MSA). Although the effectivene…

cs.CL20223 cited

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura +2

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems i…

cs.CL20211 cited

End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised Learning

Tomohiro Tanaka, Ryo Masumura, Mana Ihori +3

We propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcription-styl…

cs.CL2021

Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition

Tomohiro Tanaka, Ryo Masumura, Mana Ihori +5

We propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors. Generally,…

cs.CL2021

Unified Autoregressive Modeling for Joint End-to-End Multi-Talker Overlapped Speech Recognition and Speaker Attribute Estimation

Ryo Masumura, Daiki Okamura, Naoki Makishima +4

In this paper, we present a novel modeling method for single-channel multi-talker overlapped automatic speech recognition (ASR) systems. Fully neural network based end-to-end model…

cs.SD2021

Enrollment-less training for personalized voice activity detection

Naoki Makishima, Mana Ihori, Tomohiro Tanaka +3

We present a novel personalized voice activity detection (PVAD) learning method that does not require enrollment data during training. PVAD is a task to detect the speech segments…