97 citations · 97 across the 4 of their papers we have counts for
4 papers
CAPTDURE: Captioned Sound Dataset of Single Sources
Yuki Okamoto, Kanta Shimonishi, Keisuke Imoto +3
In conventional studies on environmental sound separation and synthesis using captions, datasets consisting of multiple-source sounds with their captions were used for model traini…
Spoofing Attacker Also Benefits from Self-Supervised Pretrained Model
Aoi Ito, Shota Horiguchi
Large-scale pretrained models using self-supervised learning have reportedly improved the performance of speech anti-spoofing. However, the attacker side may also make use of such…
Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End ASR Models
Yuki Takashima, Shota Horiguchi, Shinji Watanabe +2
In this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model. Conventional app…
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
Shinji Watanabe, Michael Mandel, Jon Barker +18
Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits…