3 citations · 3 across the 3 of their papers we have counts for
13 papers · 1 filter
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion
Alon Ziv, Harel Pogoda, Yossi Adi
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to co…
MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation
Alon Ziv, Sanyuan Chen, Andros Tjandra +3
A key challenge in music generation models is their lack of direct alignment with human preferences, as music evaluation is inherently subjective and varies widely across individua…
PAST: Phonetic-Acoustic Speech Tokenizer
Nadav Har-Tuv, Or Tal, Yossi Adi
We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike…
Salmon: A Suite for Acoustic Language Model Evaluation
Gallil Maimon, Amit Roth, Yossi Adi
Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existi…
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
Simon Rouard, Robin San Roman, Yossi Adi +1
While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that lear…
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
Ella Zeldes, Or Tal, Yossi Adi
This study introduces a refined approach to Text-to-Speech (TTS) generation that significantly enhances sampling stability across languages, with a particular focus on Hebrew. By l…