9 papers
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
Kangxiang Xia, Bingshen Mu, Xian Shi +2
Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions a…
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Kangxiang Xia, Xinfa Zhu, Jixun Yao +1
In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has pro…
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
Wenjie Tian, Xinfa Zhu, Hanke Xie +3
Recent progress in text-to-speech (TTS) has achieved impressive naturalness and flexibility, especially with the development of large language model (LLM)-based approaches. However…
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
Ziqian Wang, Zikai Liu, Xinfa Zhu +6
Generative models have excelled in audio tasks using approaches such as language models, diffusion, and flow matching. However, existing generative approaches for speech enhancemen…
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Xinsheng Wang, Mingqi Jiang, Ziyang Ma +22
Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-…
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
Boyi Kang, Xinfa Zhu, Zihan Zhang +10
Recent advancements in language models (LMs) have demonstrated strong capabilities in semantic understanding and contextual modeling, which have flourished in generative speech enh…