19 citations · 40 across the 11 of their papers we have counts for
11 papers
i-Code: An Integrative and Composable Multimodal Learning Framework
Ziyi Yang, Yuwei Fang, Chenguang Zhu +17
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to…
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition
Kenichi Kumatani, Robert Gmyr, Felipe Cruz Salinas +5
The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity. In this work, we investigate how multi-lingual Automatic Speech R…
Optimizing Alignment of Speech and Language Latent Spaces for End-to-End Speech Recognition and Understanding
Wei Wang, Shuo Ren, Yao Qian +4
The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the pe…
A Joint and Domain-Adaptive Approach to Spoken Language Understanding
Linhao Zhang, Yu Shi, Linjun Shou +3
Spoken Language Understanding (SLU) is composed of two subtasks: intent detection (ID) and slot filling (SF). There are two lines of research on SLU. One jointly tackles these two…
Transformer-F: A Transformer network with effective methods for learning universal sentence representation
Yu Shi
The Transformer model is widely used in natural language processing for sentence representation. However, the previous Transformer-based models focus on function words that have li…
Generating Human Readable Transcript for Automatic Speech Recognition with Pre-trained Language Model
Junwei Liao, Yu Shi, Ming Gong +5
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging t…