2 papers
eess.AS2026
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Eungbeom Kim, Kyogu Lee
Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments. Although distillation reduces inference…
eess.AS2026
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
Seaone Ok, Min Jun Choi, Eungbeom Kim +2
Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion me…