activity
20202024
most citedA Comparative Study of Self-supervised Speech Representation Based Voice Conversion

21 citations · 45 across the 19 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2024

Quantifying the effect of speech pathology on automatic and human speaker verification

Bence Mark Halpern, Thomas Tienkamp, Wen-Chin Huang +7

This study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verificati…

cs.SD2024

Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment

Yuka Hashizume, Li Li, Atsushi Miyashita +1

To achieve a flexible recommendation and retrieval system, it is desirable to calculate music similarity by focusing on multiple partial elements of musical pieces and allowing the…

cs.SD2023

Improving severity preservation of healthy-to-pathological voice conversion with global style tokens

Bence Mark Halpern, Wen-Chin Huang, Lester Phillip Violeta +2

In healthy-to-pathological voice conversion (H2P-VC), healthy speech is converted into pathological while preserving the identity. The paper improves on previous two-stage approach…

cs.SD2023

AAS-VC: On the Generalization Ability of Automatic Alignment Search based Non-autoregressive Sequence-to-sequence Voice Conversion

Wen-Chin Huang, Kazuhiro Kobayashi, Tomoki Toda

Non-autoregressive (non-AR) sequence-to-seqeunce (seq2seq) models for voice conversion (VC) is attractive in its ability to effectively model the temporal structure while enjoying…

cs.SD2023

Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion

Wen-Chin Huang, Tomoki Toda

Foreign accent conversion (FAC) is a special application of voice conversion (VC) which aims to convert the accented speech of a non-native speaker to a native-sounding speech with…

cs.SD202221 cited

A Comparative Study of Self-supervised Speech Representation Based Voice Conversion

Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi +1

We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attracti…