14 citations · 22 across the 7 of their papers we have counts for
12 papers
An Empirical Study on L2 Accents of Cross-lingual Text-to-Speech Systems via Vowel Space
Jihwan Lee, Jae-Sung Bae, Seongkyu Mun +4
With the recent developments in cross-lingual Text-to-Speech (TTS) systems, L2 (second-language, or foreign) accent problems arise. Moreover, running a subjective evaluation for su…
Streaming end-to-end speech recognition with jointly trained neural feature enhancement
Chanwoo Kim, Abhinav Garg, Dhananjaya Gowda +2
In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the Mo…
Overcoming label noise in audio event detection using sequential labeling
Jae-Bin Kim, Seongkyu Mun, Myungwoo Oh +3
This paper addresses the noisy label issue in audio event detection (AED) by refining strong labels as sequential labels with inaccurate timestamps removed. In AED, strong labels c…
Metric Learning for Keyword Spotting
Jaesung Huh, Minjae Lee, Heesoo Heo +2
The goal of this work is to train effective representations for keyword spotting via metric learning. Most existing works address keyword spotting as a closed-set classification pr…
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun +7
The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level repre…
The sound of my voice: speaker representation loss for target voice separation
Seongkyu Mun, Soyeon Choe, Jaesung Huh +1
Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for au…