5 papers
Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings
Byoungjun So, Jaejun Lee, Kyogu Lee
Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot ge…
EMG-to-Speech with Fewer Channels
Injune Hwang, Jaejun Lee, Kyogu Lee
Surface electromyography (EMG) is a promising modality for silent speech interfaces, but its effectiveness depends heavily on sensor placement and channel availability. In this wor…
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
Jaejun Lee, Yoori Oh, Kyogu Lee
Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in…
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
Jaejun Lee, Yoori Oh, Kyogu Lee
In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signal…
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Jaejun Lee, Kyogu Lee
In this paper, we propose Vo-Ve, a novel voice-vector embedding that captures speaker identity. Unlike conventional speaker embeddings, Vo-Ve is explainable, as it contains the pro…