2 papers
cs.SD2025
Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
Zhiyu Lin, Jingwen Yang, Jiale Zhao +3
Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approa…
cs.SD2025
Automatic Curation of Large-Scale, High-Quality, Multi-Category Music Source Separation Dataset
Ji Yu, Yang shuo, Xu Yuetonghui +3
Most current music source separation (MSS) methods rely on supervised learning, limited by training data quantity and quality. Though web-crawling can bring abundant data, platform…