2 papers
cs.SD2024
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Sijing Chen, Yuan Feng, Laipeng He +22
With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…
eess.AS2023
PP-MeT: a Real-world Personalized Prompt based Meeting Transcription System
Xiang Lyu, Yuhang Cao, Qing Wang +5
Speaker-attributed automatic speech recognition (SA-ASR) improves the accuracy and applicability of multi-speaker ASR systems in real-world scenarios by assigning speaker labels to…