3 papers
cs.SD2025
Text-Queried Audio Source Separation via Hierarchical Modeling
Xinlei Yin, Xiulian Peng, Xue Jiang +2
Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing metho…
cs.SD2025
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
Xue Jiang, Xiulian Peng, Yuan Zhang +1
Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, fol…
cs.SD2025
Latent-Domain Predictive Neural Speech Coding
Xue Jiang, Xiulian Peng, Huaying Xue +2
Neural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech c…