collaborators

6 papers

cs.SD2026

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

Gaopeng Xu, Zhenyu Wang, Zheng Xue +2

The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This pap…

eess.AS2026

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

Chengwei Liu, Shaofei Xue, Haoyin Yan +2

We propose a lightweight multi-path alignment network (LMPAN) for on-device joint acoustic echo cancellation (AEC) and noise suppression (NS) in full-duplex spoken dialogue systems…

cs.SD2026

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

Haoyin Yan, Chengwei Liu, Shaofei Xue +4

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying…

cs.SD2026

A Hybrid Discriminative and Generative System for Universal Speech Enhancement

Yinghao Liu, Chengwei Liu, Xiaotao Liang +3

Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes…

eess.AS2025

QuarkAudio Technical Report

Chengwei Liu, Haoyin Yan, Shaofei Xue +5

Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore pro…

cs.SD2025

UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens

Chengwei Liu, Haoyin Yan, Shaofei Xue +5

Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However…