works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions

Kun Zhou, You Zhang, Dianwen Ng +3

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of e…

eess.AS2025

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Alan Dao, Dinh Bach Vu, Huy Hoang Ha +6

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of spee…

eess.AS2025

ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Jiatong Shi, Jinchuan Tian, Yihan Wu +17

Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance do…

eess.AS2025

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

Haoyang Li, Jia Qi Yip, Tianyu Fan +1

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a nov…

eess.AS2024

Speech Separation using Neural Audio Codecs with Embedding Loss

Jia Qi Yip, Chin Yuen Kwok, Bin Ma +1

Neural audio codecs have revolutionized audio processing by enabling speech tasks to be performed on highly compressed representations. Recent work has shown that speech separation…