7 papers
Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean
Phannet Pov, Sovandara Chhoun, Hyun Woo Park +2
Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their training data. We study this quali…
DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition
Jaeeun Baik, Ui-Hyeop Shin, Jiwoon Lee +2
Automatic Speech Recognition (ASR) often degrades in real-world noisy environments, making noise robustness essential for deployment. Supervised noise-augmented fine-tuning is a co…
Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
Ui-Hyeop Shin, Hyung-Min Park
Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although…
Deep Filter Estimation from Inter-Frame Correlations for Monaural Speech Dereverberation
Ui-Hyeop Shin, Jun Hyung Kim, Jangyeon Kim +2
Speech dereverberation in distant-microphone scenarios remains challenging due to the high correlation between reverberation and target signals, often leading to poor generalizatio…
TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
Ui-Hyeop Shin, Bon Hyeok Ku, Hyung-Min Park
In general, multi-channel source separation has utilized inter-microphone phase differences (IPDs) concatenated with magnitude information in time-frequency domain, or real and ima…
Stack Less, Repeat More: A Block Reusing Approach for Progressive Speech Enhancement
Jangyeon Kim, Ui-Hyeop Shin, Jaehyun Ko +1
This paper presents an efficient speech enhancement (SE) approach that reuses a processing block repeatedly instead of conventional stacking. Rather than increasing the number of b…