4 papers · 1 filter
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
Or Tal, Felix Kreuk, Yossi Adi
Recent progress in text-to-music generation has enabled models to synthesize high-quality musical segments, full compositions, and even respond to fine-grained control signals, e.g…
PAST: Phonetic-Acoustic Speech Tokenizer
Nadav Har-Tuv, Or Tal, Yossi Adi
We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike…
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
Ella Zeldes, Or Tal, Yossi Adi
This study introduces a refined approach to Text-to-Speech (TTS) generation that significantly enhances sampling stability across languages, with a particular focus on Hebrew. By l…
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
Or Tal, Alon Ziv, Itai Gat +2
We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions. JASCO can generate high-quality music samples condition…