5 papers
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
Jiaqi Li, Yao Qian, Yuxuan Hu +7
Neural audio codecs are foundational to speech language models. It is expected to have a low frame rate and decoupled semantic and acoustic information. A lower frame rate codec ca…
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
Haibin Wu, Yuxuan Hu, Ruchao Fan +8
Speech language models (Speech LMs) enable end-to-end speech-text modeling within a single model, offering a promising direction for spoken dialogue systems. The choice of speech-t…
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
Yuxuan Hu, Haibin Wu, Ruchao Fan +4
Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech o…
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
Leying Zhang, Yao Qian, Xiaofei Wang +8
Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing syste…
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
Jiaqi Li, Dongmei Wang, Xiaofei Wang +13
Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the…