1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.SD2025
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
Shuyang Liu, Yuan Jin, Rui Lin +3
Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evalu…
cs.SD2025
Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
Kangdi Wang, Zhiyue Wu, Dinghao Zhou +3
Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual as…
cs.SD2025★ 1 cited
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
Xin Li, Kaikai Jia, Hao Sun +2
Recent advancements in text-to-speech (TTS) models have been driven by the integration of large language models (LLMs), enhancing semantic comprehension and improving speech natura…