activity
20162023
most citedEnglish Broadcast News Speech Recognition by Humans and Machines

12 citations · 32 across the 27 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

cs.CL2020

End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features

Edmilson Morais, Hong-Kwang J. Kuo, Samuel Thomas +2

Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in t…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…

cs.CL2020

End-to-End Spoken Language Understanding Without Full Transcripts

Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7

An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…

cs.CV2020

AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +11

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognitio…

cs.LG2020★ 1 cited

Improving Efficiency in Large-Scale Decentralized Distributed Training

Wei Zhang, Xiaodong Cui, Abdullah Kayi +9

Decentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to p…

eess.AS2020

Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard

Zoltán Tüske, George Saon, Kartik Audhkhasi +1

It is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousa…