activity
20172022
most citedEnglish Broadcast News Speech Recognition by Humans and Machines

12 citations · 16 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL2022

Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding

Vishal Sunder, Samuel Thomas, Hong-Kwang J. Kuo +3

Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in tex…

cs.CL2022

Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems

Samuel Thomas, Hong-Kwang J. Kuo, Brian Kingsbury +1

The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly proces…

cs.CL2022

Integrating Text Inputs For Training and Adapting RNN Transducer ASR Models

Samuel Thomas, Brian Kingsbury, George Saon +1

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be independently adapted to a new domain, recent end-to-en…

cs.CL2022

A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets

Zvi Kons, Aharon Satt, Hong-Kwang Kuo +4

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous…

cs.CL2022

Improving End-to-End Models for Set Prediction in Spoken Language Understanding

Hong-Kwang J. Kuo, Zoltan Tuske, Samuel Thomas +2

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts…

cs.CL2021

Cascaded Multilingual Audio-Visual Learning from Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +8

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…