19 citations · 40 across the 12 of their papers we have counts for
9 papers · 1 filter
i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data
Ziyi Yang, Mahmoud Khademi, Yichong Xu +16
The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition
Kenichi Kumatani, Robert Gmyr, Felipe Cruz Salinas +5
The sparsely-gated Mixture of Experts (MoE) can magnify a network capacity with a little computational complexity. In this work, we investigate how multi-lingual Automatic Speech R…
A Joint and Domain-Adaptive Approach to Spoken Language Understanding
Linhao Zhang, Yu Shi, Linjun Shou +3
Spoken Language Understanding (SLU) is composed of two subtasks: intent detection (ID) and slot filling (SF). There are two lines of research on SLU. One jointly tackles these two…
Transformer-F: A Transformer network with effective methods for learning universal sentence representation
Yu Shi
The Transformer model is widely used in natural language processing for sentence representation. However, the previous Transformer-based models focus on function words that have li…
Generating Human Readable Transcript for Automatic Speech Recognition with Pre-trained Language Model
Junwei Liao, Yu Shi, Ming Gong +5
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging t…
Improving Zero-shot Neural Machine Translation on Language-specific Encoders-Decoders
Junwei Liao, Yu Shi, Ming Gong +3
Recently, universal neural machine translation (NMT) with shared encoder-decoder gained good performance on zero-shot translation. Unlike universal NMT, jointly trained language-sp…