21 citations · 26 across the 10 of their papers we have counts for
10 papers
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
Injune Hwang, Kyogu Lee
Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from su…
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
Jaeyeon Kim, Injune Hwang, Kyogu Lee
We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we int…
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
Haesun Joung, Kyogu Lee
Music auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environm…
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
Jayeon Yi, Junghyun Koo, Kyogu Lee
Clipping is a common nonlinear distortion that occurs whenever the input or output of an audio system exceeds the supported range. This phenomenon undermines not only the perceptio…
Towards a New Interface for Music Listening: A User Experience Study on YouTube
Ahyeon Choi, Eunsik Shin, Haesun Joung +2
In light of the enduring success of music streaming services, it is noteworthy that an increasing number of users are positively gravitating toward YouTube as their preferred platf…
Self-refining of Pseudo Labels for Music Source Separation with Noisy Labeled Data
Junghyun Koo, Yunkee Chae, Chang-Bin Jeon +1
Music source separation (MSS) faces challenges due to the limited availability of correctly-labeled individual instrument tracks. With the push to acquire larger datasets to improv…