Publications (15)
LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers
Peidong Wang, Eric Sun, Jian Xue +5
Automatic speech recognition (ASR) and speech translation (ST) can both use neural transducers as the model structure. It is thus possible to use a single transducer model to perfo…
Building High-accuracy Multilingual ASR with Gated Language Experts and Curriculum Training
Eric Sun, Jinyu Li, Yuxuan Hu +8
We propose gated language experts and curriculum training to enhance multilingual transformer transducer models without requiring language identification (LID) input from users dur…
Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
Zhong Meng, Yu Wu, Naoyuki Kanda +6
Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition. Recently, internal language model estimat…
Self-Teaching Networks
Liang Lu, Eric Sun, Yifan Gong
We propose self-teaching networks to improve the generalization capacity of deep neural networks. The idea is to generate soft supervision labels using the output layer for trainin…
Multilingual Speech Recognition using Knowledge Transfer across Learning Processes
Rimita Lahiri, Kenichi Kumatani, Eric Sun +1
Multilingual end-to-end(E2E) models have shown a great potential in the expansion of the language coverage in the realm of automatic speech recognition(ASR). In this paper, we aim…
Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recognition
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun +7
The external language models (LM) integration remains a challenging task for end-to-end (E2E) automatic speech recognition (ASR) which has no clear division between acoustic and la…
A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability
Jian Xue, Peidong Wang, Jinyu Li +1
In this paper, we introduce our work of building a Streaming Multilingual Speech Model (SM2), which can transcribe or translate multiple spoken languages into texts of the target l…
A Configurable Multilingual Model is All You Need to Recognize All Languages
Long Zhou, Jinyu Li, Eric Sun +1
Multilingual automatic speech recognition (ASR) models have shown great promise in recent years because of the simplified model training and deployment process. Conventional method…
Reconstructing East Asian Temperatures from 1368 to 1911 Using Historical Documents, Climate Models, and Data Assimilation
Eric Sun, Kuan-hui Elaine Lin, Wan-Ling Tseng +2
We propose a novel approach for reconstructing annual temperatures in East Asia from 1368 to 1911, leveraging the Reconstructed East Asian Climate Historical Encoded Series (REACHE…
Pre-training End-to-end ASR Models with Augmented Speech Samples Queried by Text
Eric Sun, Jinyu Li, Jian Xue +1
In end-to-end automatic speech recognition system, one of the difficulties for language expansion is the limited paired speech and text training data. In this paper, we propose a n…
High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
Jinyu Li, Rui Zhao, Eric Sun +4
While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion…
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
Sunit Sivasankaran, Eric Sun, Jinyu Li +2
Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complica…
Exploring the use of AI authors and reviewers at Agents4Science
Federico Bianchi, Owen Queen, Nitya Thakkar +2
There is growing interest in using AI agents for scientific research, yet fundamental questions remain about their capabilities as scientists and reviewers. To explore these questi…
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition
Kenichi Kumatani, Robert Gmyr, Felipe Cruz Salinas +5
The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity. In this work, we investigate how multi-lingual Automatic Speech R…
Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
Zhong Meng, Naoyuki Kanda, Yashesh Gaur +6
The efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal…