3 citations · 3 across the 5 of their papers we have counts for
5 papers
Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
Younglo Lee, Shukjae Choi, Byeong-Yeol Kim +2
We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can mode…
Neural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated Full- and Sub-Band Modeling
Zhong-Qiu Wang, Samuele Cornell, Shukjae Choi +3
We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in…
Joint unsupervised and supervised learning for context-aware language identification
Jinseok Park, Hyung Yong Kim, Jihwan Park +3
Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR…
CrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis
Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju +2
While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This…
Metric Learning for User-defined Keyword Spotting
Jaemin Jung, Youkyum Kim, Jihwan Park +4
The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits t…