collaborators

6 papers

cs.LG2026

X-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

Dongjie Fu, Di Cao, Xize Cheng +6

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primar…

cs.SD2026

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

Zihan Zhang, Xize Cheng, Wenhao Yan +5

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level prec…

cs.CV2025

Open-set Cross Modal Generalization via Multimodal Unified Representation

Hai Huang, Yan Xia, Shulei Wang +6

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates…

cs.CV2025

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Hanting Wang, Tao Jin, Wang Lin +4

Bridge models in image restoration construct a diffusion process from degraded to clear images. However, existing methods typically require training a bridge model from scratch for…

cs.SD2025

T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback

Zehan Wang, Ke Lei, Chen Zhu +8

Text-to-audio (T2A) generation has achieved remarkable progress in generating a variety of audio outputs from language prompts. However, current state-of-the-art T2A models still s…

cs.CL2025

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Xize Cheng, Dongjie Fu, Xiaoda Yang +11

With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these…