5 papers
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Shuyu Li, Kejun Zhang, Jiahe Lei +5
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and…
Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
Haolin He, Renhe Sun, Zheqi Dai +16
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Au…
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Zhen Ye, Xinfa Zhu, Chi-Min Chan +17
Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and i…
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
Ziyang Jiang, Jiahe Lei, Xueyan Chen +4
Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…
Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
Zhen Ye, Peiwen Sun, Jiahe Lei +9
Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focu…