activity
20242026
collaborators

5 papers

cs.SD2026

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

Shuyu Li, Kejun Zhang, Jiahe Lei +5

Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and…

eess.AS2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

Haolin He, Renhe Sun, Zheqi Dai +16

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Au…

eess.AS2025

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Zhen Ye, Xinfa Zhu, Chi-Min Chan +17

Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and i…

cs.SD2024

pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues

Ziyang Jiang, Jiahe Lei, Xueyan Chen +4

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…

eess.AS2024

Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Zhen Ye, Peiwen Sun, Jiahe Lei +9

Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focu…