Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
Zero-Shot Text-to-Speech from Continuous Text Streams
Trung Dang, David Aponte, Dung Tran +2
Existing zero-shot text-to-speech (TTS) systems are typically designed to process complete sentences and are constrained by the maximum duration for which they have been trained. H…
cs.SD2024
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Trung Dang, David Aponte, Dung Tran +1
Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to…