4 citations · 7 across the 5 of their papers we have counts for
6 papers
An Empirical Study of Language Model Integration for Transducer based Speech Recognition
Huahuan Zheng, Keyu An, Zhijian Ou +3
Utilizing text-only data with an external language model (ELM) in end-to-end RNN-Transducer (RNN-T) for speech recognition is challenging. Recently, a class of methods such as dens…
CUSIDE: Chunking, Simulating Future Context and Decoding for Streaming ASR
Keyu An, Huahuan Zheng, Zhijian Ou +3
History and future contextual information are known to be important for accurate acoustic modeling. However, acquiring future context brings latency for streaming ASR. In this pape…
Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
Huahuan Zheng, Wenjie Peng, Zhijian Ou +1
Automatic speech recognition systems have been largely improved in the past few decades and current systems are mainly hybrid-based and end-to-end-based. The recently proposed CTC-…
Multilingual and crosslingual speech recognition using phonological-vector based phone embeddings
Chengrui Zhu, Keyu An, Huahuan Zheng +1
The use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual…
Efficient Neural Architecture Search for End-to-end Speech Recognition via Straight-Through Gradients
Huahuan Zheng, Keyu An, Zhijian Ou
Neural Architecture Search (NAS), the process of automating architecture engineering, is an appealing next step to advancing end-to-end Automatic Speech Recognition (ASR), replacin…
An empirical study of domain-agnostic semi-supervised learning via energy-based models: joint-training and pre-training
Yunfu Song, Huahuan Zheng, Zhijian Ou
A class of recent semi-supervised learning (SSL) methods heavily rely on domain-specific data augmentations. In contrast, generative SSL methods involve unsupervised learning based…