activity
20182024
most citedEspresso: A Fast End-to-end Neural Speech Recognition Toolkit

14 citations · 23 across the 4 of their papers we have counts for

collaborators

8 papers

cs.LG2024

HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR

Hainan Xu, Travis M. Bartley, Vladimir Bataev +1

We present Hybrid-Autoregressive INference TrANsducers (HAINAN), a novel architecture for speech recognition that extends the Token-and-Duration Transducer (TDT) model. Trained wit…

eess.AS2024

Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation

Nithin Rao Koluguri, Travis Bartley, Hainan Xu +4

This paper presents a new method for training sequence-to-sequence models for speech recognition and translation tasks. Instead of the traditional approach of training models on sh…

cs.SD2021

An Asynchronous WFST-Based Decoder For Automatic Speech Recognition

Hang Lv, Zhehuai Chen, Hainan Xu +3

We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech…

cs.CL201914 cited

Espresso: A Fast End-to-end Neural Speech Recognition Toolkit

Yiming Wang, Tongfei Chen, Hainan Xu +7

We present Espresso, an open-source, modular, extensible end-to-end neural automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch and the popular neu…

cs.CL20199 cited

Saliency-driven Word Alignment Interpretation for Neural Machine Translation

Shuoyang Ding, Hainan Xu, Philipp Koehn

Despite their original goal to jointly learn to align and translate, Neural Machine Translation (NMT) models, especially Transformer, are often perceived as not learning interpreta…

cs.CL2018

Improving End-to-end Speech Recognition with Pronunciation-assisted Sub-word Modeling

Hainan Xu, Shuoyang Ding, Shinji Watanabe

Most end-to-end speech recognition systems model text directly as a sequence of characters or sub-words. Current approaches to sub-word extraction only consider character sequence…