4 papers
Towards Real-Time Generative Speech Restoration with Flow-Matching
Tsun-An Hsieh, Sebastian Braun
Generative models have shown robust performance on speech enhancement and restoration tasks, but most prior approaches operate offline with high latency, making them unsuitable for…
Adaptive Deterministic Flow Matching for Target Speaker Extraction
Tsun-An Hsieh, Minje Kim
Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies…
TGIF: Talker Group-Informed Familiarization of Target Speaker Extraction
Tsun-An Hsieh, Minje Kim
State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as…
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
Tsun-An Hsieh, Heeyoul Choi, Minje Kim
Recent studies highlight the potential of textual modalities in conditioning the speech separation model's inference process. However, regularization-based methods remain underexpl…