4 papers
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Miseul Kim, Soo Jin Park, Kyungguen Byun +4
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speak…
Three Disclaimers for Safe Disclosure: A Cardwriter for Reporting the Use of Generative AI in Writing Process
Won Ik Cho, Eunjung Cho, Hyeonji Shin
Generative artificial intelligence (AI) and large language models (LLMs) are increasingly being used in the academic writing process. This is despite the current lack of unified fr…
Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting
Hyeon-Kyeong Shin, Hyewon Han, Doyeon Kim +2
In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike prev…
Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement
Doyeon Kim, Hyewon Han, Hyeon-Kyeong Shin +2
Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms…