collaborators

5 papers

cs.SD2026

A Hybrid Discriminative and Generative System for Universal Speech Enhancement

Yinghao Liu, Chengwei Liu, Xiaotao Liang +3

Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes…

eess.AS2025

Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning

YuXiang Kong, JunFeng Hou, Jian Tang +3

Large language model (LLM)-based automatic speech recognition (ASR) has recently achieved strong performance across diverse tasks, yet contextual biasing for named entities and hot…

eess.AS2025

QuarkAudio Technical Report

Chengwei Liu, Haoyin Yan, Shaofei Xue +5

Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore pro…

cs.SD2025

UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens

Chengwei Liu, Haoyin Yan, Shaofei Xue +5

Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However…

cs.CV2025

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Qijun Gan, Ruizi Yang, Jianke Zhu +2

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animati…