3 papers
cs.SD2025
MelTok: 2D Tokenization for Single-Codebook Audio Compression
Jingyi Li, Zhiyuan Zhao, Zhisheng Zhang +6
Large Audio Language Models (LALMs) have emerged with strong performance across diverse audio understanding tasks and can be further enhanced by neural audio codecs. Transitioning…
eess.AS2025
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
Yudong Yang, Zhan Liu, Wenyi Yu +3
Diffusion-based generative models have recently achieved remarkable results in speech and vocal enhancement due to their ability to model complex speech data distributions. While t…
eess.AS2025
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
Chun Yat Wu, Jiajun Deng, Guinan Li +2
Autoregressive (AR) language models have emerged as powerful solutions for zero-shot text-to-speech (TTS) synthesis, capable of generating natural speech from a few seconds of audi…