activity
20242026
most citedOn The Landscape of Spoken Language Models: A Comprehensive Survey

3 citations · 3 across the 3 of their papers we have counts for

collaborators
Showing cs.SDShow all

13 papers · 1 filter

cs.SD2026

InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

Alon Ziv, Harel Pogoda, Yossi Adi

Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to co…

cs.SD2025

MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation

Alon Ziv, Sanyuan Chen, Andros Tjandra +3

A key challenge in music generation models is their lack of direct alignment with human preferences, as music evaluation is inherently subjective and varies widely across individua…

cs.SD2025

PAST: Phonetic-Acoustic Speech Tokenizer

Nadav Har-Tuv, Or Tal, Yossi Adi

We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike…

cs.SD2025

Salmon: A Suite for Acoustic Language Model Evaluation

Gallil Maimon, Amit Roth, Yossi Adi

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existi…

cs.SD2025

MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling

Simon Rouard, Robin San Roman, Yossi Adi +1

While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that lear…

cs.SD2024

Enhancing TTS Stability in Hebrew using Discrete Semantic Units

Ella Zeldes, Or Tal, Yossi Adi

This study introduces a refined approach to Text-to-Speech (TTS) generation that significantly enhances sampling stability across languages, with a particular focus on Hebrew. By l…