papers

Publications (16)

cs.CV2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

Zhan Liu, Changli Tang, Yuxin Wang +7

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundame…

cs.SD2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Dingdong Wang, Shujie Liu, Tianhua Zhang +3

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion rec…

cs.SD2024

Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition

Guinan Li, Jiajun Deng, Youjun Chen +8

This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN s…

cs.SD2025

Effective and Efficient Mixed Precision Quantization of Speech Foundation Models

Haoning Xu, Zhaoqing Li, Zengrui Jin +7

This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estima…

cs.SD2026

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

Haoning Xu, Zhaoqing Li, Huimeng Wang +4

This paper presents a novel data-free and training-free compression approach for speech foundation models using channelwise clustering via k-means. More fine-grained, mixed sparsit…

cs.SD2026

Explainable and Trustworthy Speech Emotion Recognition Using Confidence Score and Reinforcement Learning Rectified Speech Emotion Descriptors

Youjun Chen, Xurong Xie, Mengzhe Geng +9

Explainable and trustworthy speech emotion recognition (SER) remains a challenging task to date, largely due to the scarcity of SER data with reliable speech emotion descriptor (SE…

cs.SD2025

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

Huimeng Wang, Xurong Xie, Mengzhe Geng +6

Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch a…

cs.SD2025

Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates

Haoning Xu, Zhaoqing Li, Youjun Chen +5

This paper presents a novel approach for speech foundation models compression that tightly integrates model pruning and parameter update into a single stage. Highly compact layer-l…

eess.AS2025

MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition

Chengxi Deng, Xurong Xie, Shujie Hu +10

This paper proposes a novel Mixture of Prompt-Experts based Speaker Adaptation approach (MOPSA) for elderly speech recognition. It allows zero-shot, real-time adaptation to unseen…

eess.AS2026

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

Chengxi Deng, Xurong Xie, Shujie Hu +7

This paper proposes a novel cross-utterance audio-textual prompts based speaker adaptation approach for elderly speech recognition. It enables zero-shot, real-time adaptation to un…

eess.AS2026

Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

Chengxi Deng, Xurong Xie, Shujie Hu +7

This paper proposes a novel confidence score guided incremental and speaker adaptive pseudo-labeling approach for semi-supervised elderly speech recognition. It facilitates higher-…

eess.AS2026

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

Huimeng Wang, Hui Lu, Jiajun Deng +7

Continuous autoregressive speech synthesis has recently emerged as a promising direction for zero-shot text-to-speech (TTS). However, existing methods still suffer from a fundament…

cs.SD2025

Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision

Zhaoqing Li, Haoning Xu, Zengrui Jin +7

Model compression has become an emerging need as the sizes of modern speech systems rapidly increase. In this paper, we study model weight quantization, which directly reduces the…

cs.SD2026

Covo-Audio Technical Report

Wenfu Wang, Chenxing Li, Liqiang Zhang +23

In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…

cs.SD2026

Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition

Youjun Chen, Guinan Li, Mengzhe Geng +9

This paper highlights the critical importance of multi-channel speech enhancement (MCSE) for speech emotion recognition (ER) in cocktail party scenarios. A multi-channel speech der…

cs.SD2025

Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition

Youjun Chen, Xurong Xie, Haoning Xu +6

This paper presents a novel end-to-end LLM-empowered explainable speech emotion recognition (SER) approach. Fine-grained speech emotion descriptor (SED) features, e.g., pitch, tone…