most citedA Comparative Study of Self-supervised Speech Representation Based Voice Conversion

21 citations · 26 across the 5 of their papers we have counts for

collaborators
Showing eess.ASShow all

10 papers · 1 filter

eess.AS2025

CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models

Jiajun He, Naoki Sawada, Koichi Miyazaki +1

In real-world applications, automatic speech recognition (ASR) systems must handle overlapping speech from multiple speakers and recognize rare words like technical terms. Traditio…

eess.AS2025

PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding

Jiajun He, Tomoki Toda

End-to-end automatic speech recognition (ASR) models often struggle to accurately recognize rare words. Previously, we introduced an ASR postprocessing method called error detectio…

eess.AS202414 cited

CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection

Yongyi Zang, Jiatong Shi, You Zhang +8

Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited c…

eess.AS2024

Discriminative Neighborhood Smoothing for Generative Anomalous Sound Detection

Takuya Fujimura, Keisuke Imoto, Tomoki Toda

We propose discriminative neighborhood smoothing of generative anomaly scores for anomalous sound detection. While the discriminative approach is known to achieve better performanc…

eess.AS2023

A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023

Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta +2

This paper presents our systems (denoted as T13) for the singing voice conversion challenge (SVCC) 2023. For both in-domain and cross-domain English singing voice conversion (SVC)…

eess.AS20232 cited

The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains

Erica Cooper, Wen-Chin Huang, Yu Tsao +3

We present the second edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthesized an…