38 papers
Pseudo-label distillation for discriminative anomalous sound detection
Takuya Fujimura, Tomoki Toda
Discriminative anomalous sound detection (ASD) methods train a feature extractor through a classification task using machine-information labels. They then detect anomalies in the r…
Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model
Wen-Chin Huang, Tomoki Toda
UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to pr…
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1
Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmen…
Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis
Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1
Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…
Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition
Jinyi Mi, Ding Ma, Tomoki Toda
Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language.…
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
Ding Ma, Jinyi Mi, Fengji Li +5
Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, lim…