4 papers
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
Arnon Turetzky, Avihu Dekel, Nimrod Shabtay +5
We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuo…
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
Kaizhi Qian, Xulin Fan, Junrui Ni +4
Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to cap…
A Non-autoregressive Model for Joint STT and TTS
Vishal Sunder, Brian Kingsbury, George Saon +5
In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
Slava Shechtman, Avihu Dekel
Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. V…