97 citations · 98 across the 4 of their papers we have counts for
4 papers
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
Michael Hentschel, Yuta Nishikawa, Tatsuya Komatsu +1
This study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil t…
Audio Difference Learning for Audio Captioning
Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda +1
This study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a f…
Neural Diarization with Non-autoregressive Intermediate Attractors
Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler +2
End-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single…
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
Shinji Watanabe, Michael Mandel, Jon Barker +18
Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits…