collaborators

15 papers

cs.SD2026

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

Wenxiang Guo, Changhao Pan, Ziyue Jiang +2

Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast producti…

eess.AS2026

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13

Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial rel…

eess.AS2026

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

Ke Lei, Yu Zhang, Changhao Pan +4

Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a…

cs.CV2026

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

Ruofan Hu, Menghui Zhu, Jieming Zhu +8

Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typically combine dense visual e…

cs.CL2026

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Rongsheng Zhang, Jiji Tang, Junnan Ren +6

Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn con…

eess.AS2026

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Changhao Pan, Rui Yang, Han Wang +12

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A compre…