4 papers · 1 filter
LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression
Chengwei Liu, Shaofei Xue, Haoyin Yan +2
We propose a lightweight multi-path alignment network (LMPAN) for on-device joint acoustic echo cancellation (AEC) and noise suppression (NS) in full-duplex spoken dialogue systems…
Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
Pengbo Lyu, Xiangyu Zhao, Chengwei Liu +4
We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches…
Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
YuXiang Kong, JunFeng Hou, Jian Tang +3
Large language model (LLM)-based automatic speech recognition (ASR) has recently achieved strong performance across diverse tasks, yet contextual biasing for named entities and hot…
QuarkAudio Technical Report
Chengwei Liu, Haoyin Yan, Shaofei Xue +5
Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore pro…