activity
20152021
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 60 across the 15 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2021

Cascaded Multilingual Audio-Visual Learning from Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +8

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…

cs.CL2021

Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs

Sujeong Cha, Wangrui Hou, Hyun Jung +5

A major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs wit…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…

cs.CL2019

Challenging the Boundaries of Speech Recognition: The MALACH Corpus

Michael Picheny, Zóltan Tüske, Brian Kingsbury +3

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performanc…

cs.CL20191 cited

Acoustic Model Optimization Based On Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech Recognition

Xiaodong Cui, Michael Picheny

Evolutionary stochastic gradient descent (ESGD) was proposed as a population-based approach that combines the merits of gradient-aware and gradient-free optimization algorithms for…

cs.CL201912 cited

English Broadcast News Speech Recognition by Humans and Machines

Samuel Thomas, Masayuki Suzuki, Yinghui Huang +8

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversa…