activity
20172022
most citedDNN Transfer Learning based Non-linear Feature Extraction for Acoustic Event Classification

14 citations · 22 across the 7 of their papers we have counts for

collaborators

12 papers

eess.AS2022

An Empirical Study on L2 Accents of Cross-lingual Text-to-Speech Systems via Vowel Space

Jihwan Lee, Jae-Sung Bae, Seongkyu Mun +4

With the recent developments in cross-lingual Text-to-Speech (TTS) systems, L2 (second-language, or foreign) accent problems arise. Moreover, running a subjective evaluation for su…

cs.SD2021

Streaming end-to-end speech recognition with jointly trained neural feature enhancement

Chanwoo Kim, Abhinav Garg, Dhananjaya Gowda +2

In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the Mo…

cs.SD2020

Overcoming label noise in audio event detection using sequential labeling

Jae-Bin Kim, Seongkyu Mun, Myungwoo Oh +3

This paper addresses the noisy label issue in audio event detection (AED) by refining strong labels as sequential labels with inaccurate timestamps removed. In AED, strong labels c…

eess.AS2020

Metric Learning for Keyword Spotting

Jaesung Huh, Minjae Lee, Heesoo Heo +2

The goal of this work is to train effective representations for keyword spotting via metric learning. Most existing works address keyword spotting as a closed-set classification pr…

eess.AS2020

In defence of metric learning for speaker recognition

Joon Son Chung, Jaesung Huh, Seongkyu Mun +7

The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level repre…

cs.SD20192 cited

The sound of my voice: speaker representation loss for target voice separation

Seongkyu Mun, Soyeon Choe, Jaesung Huh +1

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for au…