collaborators

11 papers

cs.SD2026

MMAG: A Multi-Control Mixed Audio Generation Benchmark

Zihao Zheng, Xuenan Xu, Jiahao Mei +5

Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…

cs.IR2026

TWICE: Modeling the Temporal Evolution of Personalized User Behavior via Event-Driven Agents

Bingrui Jin, Kunyao Lan, Baihan LI +1

User simulators are widely used for data generation, evaluation, and agent-based interaction, but existing approaches often model users as static personas or rely on generic histor…

cs.SD2026

CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS

Zihao Zheng, Wen Wu, Chao Zhang +2

Current Text-to-Speech (TTS) systems typically use separate models for speech-prompted and text-prompted timbre control. While unifying both control signals into a single model is…

cs.SD2026

LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment

Jiahao Mei, Xuenan Xu, Zeyu Xie +4

Recent advances in text-to-music models have enabled coherent music generation from text prompts, yet fine-grained emotional control remains unresolved. We introduce LARA-Gen, a fr…

cs.SD2026

SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents

Zeyu Xie, Chenxing Li, Qiao Jin +6

Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstr…

cs.SD2026

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Ye Tao, Wen Wu, Chao Zhang +3

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…