4 papers
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
Haoyuan Yang, Mu Yang, Jiamin Xie +2
Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacit…
Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
Mu Yang, John H. L. Hansen
Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains cha…
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
Mu Yang, Szu-Jui Chen, Jiamin Xie +1
One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete token-based para…
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
Mu Yang, Bowen Shi, Matthew Le +2
This work focuses on improving Text-To-Audio (TTA) generation on zero-shot and few-shot settings (i.e. generating unseen or uncommon audio events). Inspired by the success of Retri…