10 papers
Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity
Zhenwei Mou, Liping Chen, Yajun Hu +3
Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based…
An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization
Xiao-Hang Jiang, Yang Ai, Fei Liu +4
Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec,…
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
Jing-Xuan Zhang, Genshun Wan, Jin Li +3
While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…
Rethinking Practical and Efficient Quantization Calibration for Vision-Language Models
Zhenhao Shang, Haizhao Jing, Guoting Wei +4
Post-training quantization (PTQ) is a primary approach for deploying large language models without fine-tuning, and the quantized performance is often strongly affected by the cali…
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
Genshun Wan, Wenhui Zhang, Jing-Xuan Zhang +3
Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within t…
SLM-SS: Speech Language Model for Generative Speech Separation
Tianhua Li, Chenda Li, Wei Wang +4
Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to…