collaborators

38 papers

eess.AS2026

Pseudo-label distillation for discriminative anomalous sound detection

Takuya Fujimura, Tomoki Toda

Discriminative anomalous sound detection (ASD) methods train a feature extractor through a classification task using machine-information labels. They then detect anomalies in the r…

cs.SD2026

Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

Wen-Chin Huang, Tomoki Toda

UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to pr…

cs.SD2026

PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment

Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1

Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmen…

cs.SD2026

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…

cs.SD2026

Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition

Jinyi Mi, Ding Ma, Tomoki Toda

Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language.…

eess.AS2026

Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

Ding Ma, Jinyi Mi, Fengji Li +5

Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, lim…