activity
20172026
most citedLong-span language modeling for speech recognition

8 citations · 20 across the 47 of their papers we have counts for

collaborators
Showing cs.SDShow all

24 papers · 1 filter

cs.SD2026

FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining

Xiquan Li, Xuenan Xu, Ziyang Ma +4

Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying g…

cs.SD2026

Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition

Yuhang Dai, Haopeng Lin, Jiale Qian +9

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limi…

cs.SD2026

Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models

Xiquan Li, Junxi Liu, Wenxi Chen +3

Reinforcement Learning (RL) has become an effective paradigm for enhancing Large Language Models (LLMs) and visual generative models. However, its application in text-to-audio (TTA…

cs.SD2026

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

Ziyang Ma, Ruiyang Xu, Yinghao Ma +9

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Cha…

cs.SD2026

The SJTU X-LANCE Lab System for MSR Challenge 2025

Jinxuan Zhu, Hao Qiu, Haina Zhu +3

This report describes the system submitted to the music source restoration (MSR) Challenge 2025. Our approach is composed of sequential BS-RoFormers, each dealing with a single tas…

cs.SD2026

Audio ControlNet for Fine-Grained Audio Generation and Editing

Haina Zhu, Yao Xiao, Xiquan Li +5

We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over at…