Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Eungbeom Kim, Kyogu Lee
Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments. Although distillation reduces inference…
eess.AS2026
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
Seaone Ok, Min Jun Choi, Eungbeom Kim +2
Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion me…
eess.AS2024
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
Eungbeom Kim, Hantae Kim, Kyogu Lee
Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR d…