5 citations · 5 across the 5 of their papers we have counts for
5 papers
Transduce and Speak: Neural Transducer for Text-to-Speech with Semantic Token Prediction
Minchan Kim, Myeonghun Jeong, Byoung Jin Choi +2
We introduce a text-to-speech(TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec2.0 embeddings, which makes it easy to adopt a ne…
Towards single integrated spoofing-aware speaker verification embeddings
Sung Hwan Mun, Hye-jin Shim, Hemlata Tak +12
This study aims to develop a single integrated spoofing-aware speaker verification (SASV) embeddings that satisfy two aspects. First, rejecting non-target speakers' input as well a…
When Crowd Meets Persona: Creating a Large-Scale Open-Domain Persona Dialogue Corpus
Won Ik Cho, Yoon Kyung Lee, Seoyeon Bae +5
Building a natural language dataset requires caution since word semantics is vulnerable to subtle text change or the definition of the annotated concept. Such a tendency can be see…
Disentangled Speaker Representation Learning via Mutual Information Minimization
Sung Hwan Mun, Min Hyun Han, Minchan Kim +2
Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unrave…
Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
Sung Hwan Mun, Min Hyun Han, Dongjune Lee +2
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and…