1 paper
Christophe Van Gysel, Maggie Wu, Lyan Verwimp +4
End-to-end (E2E) Automatic Speech Recognition (ASR) models are trained using paired audio-text samples that are expensive to obtain, since high-quality ground-truth data requires h…