activity
20242026
collaborators

12 papers

cs.SD2026

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou +8

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…

cs.SD2026

AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning

Yan Rong, Fengji Ma, Xu Li +3

Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle…

cs.SD2026

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Jiaqi Li, Chaoren Wang, Xiaohai Tian +9

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying informati…

cs.SD2026

SegTune: Structured and Fine-Grained Control for Song Generation

Yuejiao Wang, Zihao Ji, Pengfei Cai +6

Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attribu…

cs.SD2026

SegTune: Structured and Fine-Grained Control for Song Generation

Pengfei Cai, Joanna Wang, Haorui Zheng +6

Recent advancements in song generation have shown promising results in generating songs from lyrics and/or global text prompts. However, most existing systems lack the ability to m…

eess.AS2026

TASU: Text-Only Alignment for Speech Understanding

Jing Peng, Yi Yang, Xu Li +5

Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks. However, prevailing alignment…