12 citations · 32 across the 27 of their papers we have counts for
6 papers · 1 filter
End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features
Edmilson Morais, Hong-Kwang J. Kuo, Samuel Thomas +2
Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in t…
Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems
Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5
Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…
End-to-End Spoken Language Understanding Without Full Transcripts
Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7
An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…
AVLnet: Learning Audio-Visual Language Representations from Instructional Videos
Andrew Rouditchenko, Angie Boggust, David Harwath +11
Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognitio…
Improving Efficiency in Large-Scale Decentralized Distributed Training
Wei Zhang, Xiaodong Cui, Abdullah Kayi +9
Decentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to p…
Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard
Zoltán Tüske, George Saon, Kartik Audhkhasi +1
It is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousa…