34 citations · 46 across the 3 of their papers we have counts for
4 papers
Towards Realistic Visual Dubbing with Heterogeneous Sources
Tianyi Xie, Liucheng Liao, Cheng Bi +7
The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current appro…
Improving RNN transducer with normalized jointer network
Mingkun Huang, Jun Zhang, Meng Cai +5
Recurrent neural transducer (RNN-T) is a promising end-to-end (E2E) model in automatic speech recognition (ASR). It has shown superior performance compared to traditional hybrid AS…
Universal Phone Recognition with a Multilingual Allophone System
Xinjian Li, Siddharth Dalmia, Juncheng Li +8
Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, genera…
Real-time Neural-based Input Method
Jiali Yao, Raphael Shu, Xinjian Li +2
The input method is an essential service on every mobile and desktop devices that provides text suggestions. It converts sequential keyboard inputs to the characters in its target…