4 papers
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ã-Tsang, Amdo and Kham Speech Dataset Generation
Yutong Liu, Ziyue Zhang, Ban Ma-bao +7
Tibetan is a low-resource language with minimal parallel speech corpora spanning its three major dialects-Ã-Tsang, Amdo, and Kham-limiting progress in speech modeling. To address…
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ã-Tsang, Amdo and Kham Speech Dataset Generation
Yutong Liu, Ziyue Zhang, Ban Ma-bao +7
Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ã-Tsang, Amdo, and Kham), limiting progress in speech modeling. To addre…
Listening, Imagining & Refining: A Heuristic Optimized ASR Correction Framework with LLMs
Yutong Liu, Ziyue Zhang, Cheng Huang +4
Automatic Speech Recognition (ASR) systems remain prone to errors that affect downstream applications. In this paper, we propose LIR-ASR, a heuristic optimized iterative correction…
TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
Yutong Liu, Feng Xiao, Ziyue Zhang +10
Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction…