activity
20172023
most citedGated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion

11 citations · 24 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2023

Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding

Suyoun Kim, Akshat Shrivastava, Duc Le +3

End-to-end (E2E) spoken language understanding (SLU) systems that generate a semantic parse from speech have become more promising recently. This approach uses a single model that…

cs.CL2022

Introducing Semantics into Speech Encoders

Derek Xu, Shuyan Dong, Changhan Wang +10

Recent studies find existing self-supervised speech encoders contain primarily acoustic rather than semantic information. As a result, pipelined supervised automatic speech recogni…

cs.CL2022

Joint Audio/Text Training for Transformer Rescorer of Streaming Speech Recognition

Suyoun Kim, Ke Li, Lucas Kabela +4

Recently, there has been an increasing interest in two-pass streaming end-to-end speech recognition (ASR) that incorporates a 2nd-pass rescoring model on top of the conventional 1s…

cs.CL20215 cited

Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding

Suyoun Kim, Abhinav Arora, Duc Le +4

Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator…

cs.CL2021

Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion

Duc Le, Mahaveer Jain, Gil Keren +9

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for sp…

cs.CL20201 cited

Improving RNN Transducer Based ASR with Auxiliary Tasks

Chunxi Liu, Frank Zhang, Duc Le +3

End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recogni…