40 citations · 199 across the 19 of their papers we have counts for
7 papers · 1 filter
Serial-OE: Anomalous sound detection based on serial method with outlier exposure capable of using small amounts of anomalous data for training
Ibuki Kuroyanagi, Tomoki Hayashi, Kazuya Takeda +1
We introduce Serial-OE, a new approach to anomalous sound detection (ASD) that leverages small amounts of anomalous data to improve the performance. Conventional ASD methods rely p…
ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit
Brian Yan, Jiatong Shi, Yun Tang +13
ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community. ESPnet-ST-v2 supports 1) offlin…
Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem
Jing Shi, Xuankai Chang, Tomoki Hayashi +3
Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain a…
S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations
Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi +3
This paper introduces S3PRL-VC, an open-source voice conversion (VC) framework based on the S3PRL toolkit. In the context of recognition-synthesis VC, self-supervised speech repres…
On Prosody Modeling for ASR+TTS based Voice Conversion
Wen-Chin Huang, Tomoki Hayashi, Xinjian Li +2
In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to t…
Anomalous Sound Detection Using a Binary Classification Model and Class Centroids
Ibuki Kuroyanagi, Tomoki Hayashi, Kazuya Takeda +1
An anomalous sound detection system to detect unknown anomalous sounds usually needs to be built using only normal sound data. Moreover, it is desirable to improve the system by ef…