activity
20242026
collaborators

13 papers

cs.SD2026

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

Xiao-Hang Jiang, Han-Jie Guo, Ying-Si Liang +4

Ambient clinical scribes increasingly combine Automatic Speech Recognition with Large Language Models to automate documentation. However, traditional metrics like Word Error Rate m…

eess.AS2026

VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization

Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng +3

Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits…

eess.AS2026

An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization

Xiao-Hang Jiang, Yang Ai, Fei Liu +4

Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec,…

eess.AS2026

CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement

Xiao-Hang Jiang, Yang Ai, Hui-Peng Du +2

High-quality speech coding at low bitrates is crucial for bandwidth-constrained applications, yet remains challenging due to the severe loss of quality-critical information in high…

eess.AS2026

Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction

Hui-Peng Du, Yang Ai, Xiao-Hang Jiang +2

Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budget…

eess.AS2026

CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction

Hui-Peng Du, Yang Ai, Xiao-Hang Jiang +2

This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for ef…