4 papers
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition
Kunyuan Xie, Zhixi Cai, Kalin Stefanov
Subtle hand differences make sign language recognition challenging, yet many existing methods rely on encoders pretrained on generic action datasets that poorly capture such fine-g…
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Kun Xie, Feiyu Shen, Junjie Li +3
Current dialogue generation approaches typically require the complete dialogue text before synthesis and produce a single, inseparable speech containing all voices, making them uns…
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
Hao-Han Guo, Yao Hu, Fei-Yu Shen +4
In this work, we upgrade FireRedTTS to a new version, FireRedTTS-1S, a high-quality streaming foundation text-to-speech system. FireRedTTS-1S achieves streaming speech generation v…
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Hao-Han Guo, Yao Hu, Kun Liu +6
This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications. The framework compris…