9 citations · 18 across the 3 of their papers we have counts for
4 papers
End-to-End Multi-Channel Transformer for Speech Recognition
Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris +2
Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectur…
Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding
Bhuvan Agrawal, Markus Müller, Martin Radfar +3
End-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a…
End-to-End Neural Transformer Based Spoken Language Understanding
Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann
Spoken language understanding (SLU) refers to the process of inferring the semantic information from audio signals. While the neural transformers consistently deliver the best perf…
Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy +11
Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language ide…