collaborators

8 papers

cs.SD2026

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou +8

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…

cs.CL2026

TriggerBench: Investigating Prospective Memory for Large Language Models

Tianhua Zhang, Xinjiang Wang, Qianxi Zhang +6

While Large Language Models (LLMs) are increasingly deployed in long interactions, existing evaluations focus predominantly on retrospective memory (RM) via explicit queries. Prosp…

cs.AI2026

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Jialiang Yang, Bin Xia, Ruihang Chu +6

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactio…

cs.SD2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Dingdong Wang, Shujie Liu, Tianhua Zhang +3

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion rec…

cs.CL2026

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Dingdong Wang, Junan Li, Jincenzi Wu +4

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requi…

cs.SD2026

V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation

Nolan Chan, Timmy Gang, Yongqian Wang +2

This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations…