4 papers
ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure
Zixiang Wan, Xusheng Yang, Zheng Wang +1
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to pr…
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
Zixiang Wan, Haoran Zhao, Guochang Zhang +3
This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an opti…
SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
Zixiang Wan, Guochang Zhang, Yifeng He +1
Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream…
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
Zixiang Wan, Ziyue Qiu, Yiyang Liu +1
Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio inform…