14 papers
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
Hao Zhang, Weiwei Li, Rilin Chen +3
Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice…
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang +23
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…
Scalable Neural Vocoder from Range-Null Space Decomposition
Andong Li, Tong Lei, Zhihang Sun +4
Although deep neural networks have facilitated significant progress of neural vocoders in recent years, they usually suffer from intrinsic challenges like opaque modeling, inflexib…
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
Junqi Zhao, Chenxing Li, Jinzheng Zhao +4
We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis…
Gen-SER: When the generative model meets speech emotion recognition
Taihui Wang, Jinzheng Zhao, Rilin Chen +3
Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different fr…
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
Andong Li, Tong Lei, Rilin Chen +5
This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare…