activity
20192025
most citedSupervised and Unsupervised Approaches for Controlling Narrow Lexical Focus in Sequence-to-Sequence Speech Synthesis

9 citations · 11 across the 4 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2024

Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer

Slava Shechtman, Avihu Dekel

Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. V…

eess.AS2024

Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion

Arnon Turetzky, Avihu Dekel, Nimrod Shabtay +5

We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuo…

eess.AS20219 cited

Supervised and Unsupervised Approaches for Controlling Narrow Lexical Focus in Sequence-to-Sequence Speech Synthesis

Slava Shechtman, Raul Fernandez, David Haws

Although Sequence-to-Sequence (S2S) architectures have become state-of-the-art in speech synthesis, capable of generating outputs that approach the perceptual quality of natural sa…

eess.AS20202 cited

Controllable Sequence-To-Sequence Neural TTS with LPCNET Backend for Real-time Speech Synthesis on CPU

Slava Shechtman, Carmel Rabinovitz, Alex Sorin +2

State-of-the-art sequence-to-sequence acoustic networks, that convert a phonetic sequence to a sequence of spectral features with no explicit prosody prediction, generate speech wi…

eess.AS2019

Sequence to Sequence Neural Speech Synthesis with Prosody Modification Capabilities

Slava Shechtman, Alex Sorin

Modern sequence to sequence neural TTS systems provide close to natural speech quality. Such systems usually comprise a network converting linguistic/phonetic features sequence to…

eess.AS2019

High quality, lightweight and adaptable TTS using LPCNet

Zvi Kons, Slava Shechtman, Alex Sorin +2

We present a lightweight adaptable neural TTS system with high quality output. The system is composed of three separate neural network blocks: prosody prediction, acoustic feature…