21 citations · 45 across the 19 of their papers we have counts for
8 papers · 1 filter
Quantifying the effect of speech pathology on automatic and human speaker verification
Bence Mark Halpern, Thomas Tienkamp, Wen-Chin Huang +7
This study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verificati…
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
Yuka Hashizume, Li Li, Atsushi Miyashita +1
To achieve a flexible recommendation and retrieval system, it is desirable to calculate music similarity by focusing on multiple partial elements of musical pieces and allowing the…
Improving severity preservation of healthy-to-pathological voice conversion with global style tokens
Bence Mark Halpern, Wen-Chin Huang, Lester Phillip Violeta +2
In healthy-to-pathological voice conversion (H2P-VC), healthy speech is converted into pathological while preserving the identity. The paper improves on previous two-stage approach…
AAS-VC: On the Generalization Ability of Automatic Alignment Search based Non-autoregressive Sequence-to-sequence Voice Conversion
Wen-Chin Huang, Kazuhiro Kobayashi, Tomoki Toda
Non-autoregressive (non-AR) sequence-to-seqeunce (seq2seq) models for voice conversion (VC) is attractive in its ability to effectively model the temporal structure while enjoying…
Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion
Wen-Chin Huang, Tomoki Toda
Foreign accent conversion (FAC) is a special application of voice conversion (VC) which aims to convert the accented speech of a non-native speaker to a native-sounding speech with…
A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi +1
We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attracti…