collaborators

5 papers

cs.SD2026

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

Wei Zhou, Wanyi Ning, Yinshang Guo +3

Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic co…

cs.SD2026

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Wanyi Ning, Wei Zhou, Yingpeng Li +3

Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision ar…

cs.CL2026

FormalASR: End-to-End Spoken Chinese to Formal Text

Wanyi Ning, Yinshang Guo, Haitao Qian +3

Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are o…

cs.CL2026

EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer

Hao Zhang, Zhibin Zhang, Guangxin Wu +3

Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinder deployment in resource-con…

cs.LG2025

MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration

Jinguang Wang, Jingyu Wang, Haifeng Sun +6

Quantization has been widely used to compress and accelerate inference of large language models (LLMs). Existing methods focus on exploring the per-token dynamic calibration to ens…