4 papers · 1 filter
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
Zhicheng Ouyang, Seong-Gyun Leem, Bach Viet Do +4
Conversational AI has made significant progress, yet generating expressive and controllable text-to-speech (TTS) remains challenging. Specifically, controlling fine-grained voice s…
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
Biel Tura Vecino, Subhadeep Maji, Aravind Varier +11
The use of audio recordings of human speech to train LLMs poses privacy concerns due to these models' potential to generate outputs that closely resemble artifacts in the training…
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
Yen-Ju Lu, Jing Liu, Thomas Thebaud +4
We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to s…
Speech Recognition Rescoring with Large Speech-Text Foundation Models
Prashanth Gurunath Shivakumar, Jari Kolehmainen, Aditya Gourav +4
Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often…