12 papers
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
Yunrui Cai, Xu Li, Yucheng Zhou +8
Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…
AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning
Yan Rong, Fengji Ma, Xu Li +3
Time-aware dense audio captioning (TDAC) aims to generate multiple fine-grained attributes (dense) of the audio with precise time boundaries (time-aware). Existing methods struggle…
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model
Jiaqi Li, Chaoren Wang, Xiaohai Tian +9
Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying informati…
SegTune: Structured and Fine-Grained Control for Song Generation
Yuejiao Wang, Zihao Ji, Pengfei Cai +6
Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attribu…
SegTune: Structured and Fine-Grained Control for Song Generation
Pengfei Cai, Joanna Wang, Haorui Zheng +6
Recent advancements in song generation have shown promising results in generating songs from lyrics and/or global text prompts. However, most existing systems lack the ability to m…
TASU: Text-Only Alignment for Speech Understanding
Jing Peng, Yi Yang, Xu Li +5
Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks. However, prevailing alignment…