9 citations · 22 across the 4 of their papers we have counts for
3 papers · 1 filter
Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding
Bhuvan Agrawal, Markus Müller, Martin Radfar +3
End-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a…
Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy +11
Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language ide…
Phonemic and Graphemic Multilingual CTC Based Speech Recognition
Markus Müller, Sebastian Stüker, Alex Waibel
Training automatic speech recognition (ASR) systems requires large amounts of data in the target language in order to achieve good performance. Whereas large training corpora are r…