8 papers
Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training
Wonchul Shin, Inyong Choi, Kyogu Lee
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis rev…
Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings
Byoungjun So, Jaejun Lee, Kyogu Lee
Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot ge…
EMG-to-Speech with Fewer Channels
Injune Hwang, Jaejun Lee, Kyogu Lee
Surface electromyography (EMG) is a promising modality for silent speech interfaces, but its effectiveness depends heavily on sensor placement and channel availability. In this wor…
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
Jaejun Lee, Yoori Oh, Kyogu Lee
Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in…
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
Jaejun Lee, Yoori Oh, Kyogu Lee
In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signal…
DOSE : Drum One-Shot Extraction from Music Mixture
Suntae Hwang, Seonghyeon Kang, Kyungsu Kim +2
Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal…