activity
20242026
collaborators

22 papers

cs.SD2026

FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

Shiyao Wang, Xijuan Zeng, Hui Wang +4

We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized,…

cs.SD2026

SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

Hui Wang, Jinghua Zhao, Yifan Yang +9

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scala…

cs.SD2026

GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR

Yujie Guo, Jiaming Zhou, Yuhang Jia +2

End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech. A critical bottleneck is that speaker-speci…

cs.SD2026

Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models

Haoqin Sun, Chenyang Lyu, Shiwan Zhao +5

Despite the growing success of Large Speech Language Models (LSLMs) in processing short-term acoustic signals, their extension to long-form audio understanding is severely bottlene…

cs.SD2026

DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding

Jiaming Zhou, Xuxin Cheng, Shiwan Zhao +5

Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains cost…

cs.SD2025

DIFFA: Large Language Diffusion Models Can Listen and Understand

Jiaming Zhou, Hongjie Chen, Shiwan Zhao +9

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged…