21 papers
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
Doyeop Kwak, Youngjoon Jang, Seongyu Kim +1
Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individua…
FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee +1
In this work, we present FastAV, the first token pruning framework tailored for audio-visual large language models (AV-LLMs). While token pruning has been actively explored in stan…
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
Jihoo Jung, Ji-Hoon Kim, Doyeop Kwak +3
We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly…
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
Jeongsoo Choi, Jaehun Kim, Joon Son Chung
This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, an…
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
Doyeop Kwak, Youngjoon Jang, Joon Son Chung
The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional genera…
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak +2
The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conv…