collaborators

9 papers

eess.AS2026

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

Chengwei Liu, Shaofei Xue, Haoyin Yan +2

We propose a lightweight multi-path alignment network (LMPAN) for on-device joint acoustic echo cancellation (AEC) and noise suppression (NS) in full-duplex spoken dialogue systems…

cs.SD2026

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

Haoyin Yan, Chengwei Liu, Shaofei Xue +4

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying…

eess.AS2026

Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models

Pengbo Lyu, Xiangyu Zhao, Chengwei Liu +4

We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches…

cs.SD2026

A Hybrid Discriminative and Generative System for Universal Speech Enhancement

Yinghao Liu, Chengwei Liu, Xiaotao Liang +3

Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes…

eess.AS2025

Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning

YuXiang Kong, JunFeng Hou, Jian Tang +3

Large language model (LLM)-based automatic speech recognition (ASR) has recently achieved strong performance across diverse tasks, yet contextual biasing for named entities and hot…

eess.AS2025

QuarkAudio Technical Report

Chengwei Liu, Haoyin Yan, Shaofei Xue +5

Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore pro…