9 citations · 20 across the 7 of their papers we have counts for
3 papers · 1 filter
End-to-End Multi-Channel Transformer for Speech Recognition
Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris +2
Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectur…
Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding
Bhuvan Agrawal, Markus Müller, Martin Radfar +3
End-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a…
Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy +11
Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language ide…