activity
20172025
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 30 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2025

LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

Rao Ma, Tongzhou Chen, Kartik Audhkhasi +1

Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language process…

cs.CL2024

STAB: Speech Tokenizer Assessment Benchmark

Shikhar Vashishth, Harman Singh, Shikhar Bharadwaj +6

Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the wi…

cs.CL2022

Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition

Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran +1

Attention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture. Attention is typ…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…

cs.CL2020

End-to-End Spoken Language Understanding Without Full Transcripts

Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7

An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…

cs.CL2019

Challenging the Boundaries of Speech Recognition: The MALACH Corpus

Michael Picheny, Zóltan Tüske, Brian Kingsbury +3

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performanc…