activity
20242026
most citedWavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convolution and Harmonic Prior for Reliable Complex Spectrogram Estimation

1 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing cs.SDShow all

27 papers · 1 filter

cs.SD2026

Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

Wen-Chin Huang, Tomoki Toda

UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to pr…

cs.SD2026

PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment

Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1

Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmen…

cs.SD2026

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

Masato Murata, Koichi Miyazaki, Tomoki Koriyama +1

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory durin…

cs.SD2026

Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition

Jinyi Mi, Ding Ma, Tomoki Toda

Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language.…

cs.SD2026

An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results

Lester Phillip Violeta, Xueyao Zhang, Jiatong Shi +4

We present a thorough analysis of the findings of the latest iteration of the Singing Voice Conversion Challenge, a scientific event aiming to compare and understand different voic…

cs.SD2026

MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models

Wen-Chin Huang, Erica Cooper, Tomoki Toda

In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neura…