1 citations · 1 across the 10 of their papers we have counts for
17 papers
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
Haoqin Sun, Chenyang Lyu, Shiwan Zhao +5
Despite the growing success of Large Speech Language Models (LSLMs) in processing short-term acoustic signals, their extension to long-form audio understanding is severely bottlene…
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
Jiaming Zhou, Xuxin Cheng, Shiwan Zhao +5
Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains cost…
Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
Shiyao Wang, Shiwan Zhao, Jiaming Zhou +1
Dysarthric speech recognition (DSR) research has witnessed remarkable progress in recent years, evolving from the basic understanding of individual words to the intricate comprehen…
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
Haoqin Sun, Chenyang Lyu, Xiangyu Kong +9
Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discr…
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
Hui Wang, Cheng Liu, Junyang Chen +7
Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generaliza…
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
Enzhi Wang, Qicheng Li, Shiwan Zhao +6
In recent years, large language models (LLMs) have achieved remarkable advancements in multimodal processing, including end-to-end speech-based language models that enable natural…