6 citations · 10 across the 18 of their papers we have counts for
6 papers · 1 filter
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
Doyeop Kwak, Youngjoon Jang, Joon Son Chung
The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional genera…
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
Doyeop Kwak, Youngjoon Jang, Seongyu Kim +1
Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individua…
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang +2
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key componen…
Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model
Suyeon Lee, Chaeyoung Jung, Youngjoon Jang +2
The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their…
Metric Learning for User-defined Keyword Spotting
Jaemin Jung, Youkyum Kim, Jihwan Park +4
The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits t…