collaborators

20 papers

cs.SD2026

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Jinting Wang, Yuguang Yang, Shengyu Li +4

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated…

eess.AS2026

Audio Editing in the Era of Foundation Models: A Survey

Changhao Pan, Yifei Fan, Fan Zhuo +13

Audio editing aims to modify a given synthetic or real-world audio signal to satisfy specific user needs. As a promising yet challenging direction in AIGC, it has attracted increas…

eess.AS2026

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

Jingyu Lu, Yuhan Wang, Jianming Luo +15

More than a dozen spoken dialogue systems have recently claimed to be "full-duplex," yet the term has been used to describe substantially different capabilities. Existing surveys c…

cs.CV2026

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

Ruofan Hu, Menghui Zhu, Jieming Zhu +8

Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typically combine dense visual e…

cs.SD2026

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

Xiaoda Yang, Majun Zhang, Changhao Pan +10

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from g…

cs.CV2026

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

Jiayang Xu, Fan Zhuo, Majun Zhang +6

Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoup…