1 paper
Li Fu, Xiaoxiao Li, Runyu Wang +5
End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervisio…