activity
20172022
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 30 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CL2022

Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition

Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran +1

Attention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture. Attention is typ…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…

cs.CL2020

End-to-End Spoken Language Understanding Without Full Transcripts

Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7

An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…

eess.AS2020

Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard

Zoltán Tüske, George Saon, Kartik Audhkhasi +1

It is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousa…

cs.CL2019

Challenging the Boundaries of Speech Recognition: The MALACH Corpus

Michael Picheny, Zóltan Tüske, Brian Kingsbury +3

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performanc…

cs.CL20193 cited

Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation

Gakuto Kurata, Kartik Audhkhasi

Conventional automatic speech recognition (ASR) systems trained from frame-level alignments can easily leverage posterior fusion to improve ASR accuracy and build a better single m…