34 citations · 107 across the 13 of their papers we have counts for
16 papers
Shallow Fusion of Weighted Finite-State Transducer and Language Model for Text Normalization
Evelina Bakhturina, Yang Zhang, Boris Ginsburg
Text normalization (TN) systems in production are largely rule-based using weighted finite-state transducers (WFST). However, WFST-based systems struggle with ambiguous input when…
Towards Realistic Visual Dubbing with Heterogeneous Sources
Tianyi Xie, Liucheng Liao, Cheng Bi +7
The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current appro…
Revisiting IPA-based Cross-lingual Text-to-speech
Haitong Zhang, Haoyue Zhan, Yang Zhang +2
International Phonetic Alphabet (IPA) has been widely used in cross-lingual text-to-speech (TTS) to achieve cross-lingual voice cloning (CL VC). However, IPA itself has been unders…
A Unified Transformer-based Framework for Duplex Text Normalization
Tuan Manh Lai, Yang Zhang, Evelina Bakhturina +2
Text normalization (TN) and inverse text normalization (ITN) are essential preprocessing and postprocessing steps for text-to-speech synthesis and automatic speech recognition, res…
SGD-QA: Fast Schema-Guided Dialogue State Tracking for Unseen Services
Yang Zhang, Vahid Noroozi, Evelina Bakhturina +1
Dialogue state tracking is an essential part of goal-oriented dialogue systems, while most of these state tracking models often fail to handle unseen services. In this paper, we pr…
NeMo Inverse Text Normalization: From Development To Production
Yang Zhang, Evelina Bakhturina, Kyle Gorman +1
Inverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text to improve the readability of the ASR output. Many state-…