3 papers
cs.SD2025
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
Or Tal, Felix Kreuk, Yossi Adi
Recent progress in text-to-music generation has enabled models to synthesize high-quality musical segments, full compositions, and even respond to fine-grained control signals, e.g…
cs.SD2025
PAST: Phonetic-Acoustic Speech Tokenizer
Nadav Har-Tuv, Or Tal, Yossi Adi
We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike…
cs.SD2024
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
Ella Zeldes, Or Tal, Yossi Adi
This study introduces a refined approach to Text-to-Speech (TTS) generation that significantly enhances sampling stability across languages, with a particular focus on Hebrew. By l…