5 citations · 5 across the 3 of their papers we have counts for
3 papers · 1 filter
Transduce and Speak: Neural Transducer for Text-to-Speech with Semantic Token Prediction
Minchan Kim, Myeonghun Jeong, Byoung Jin Choi +2
We introduce a text-to-speech(TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec2.0 embeddings, which makes it easy to adopt a ne…
Disentangled Speaker Representation Learning via Mutual Information Minimization
Sung Hwan Mun, Min Hyun Han, Minchan Kim +2
Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unrave…
Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
Sung Hwan Mun, Min Hyun Han, Dongjune Lee +2
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and…