collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

Jiabao Zhuang, Changhao Jiang, Hanchen Wang +11

Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligni…

cs.SD2026

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

Hui Li, Yangfan Gao, Junlin Shang +4

Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and…

cs.SD2026

Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control

Changhao Jiang, Jiahao Chen, Zhenghao Xiang +14

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of pub…

cs.SD2026

CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges

Hui Li, Changhao Jiang, Hongyu Wang +11

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks…

cs.SD2025

From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling

Yifei Cao, Changhao Jiang, Jiabao Zhuang +11

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on huma…