3 papers
eess.AS2025
SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
Zixiang Wan, Guochang Zhang, Yifeng He +1
Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream…
eess.AS2025
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
Zixiang Wan, Haoran Zhao, Guochang Zhang +3
This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an opti…
eess.AS2024
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
Zixiang Wan, Ziyue Qiu, Yiyang Liu +1
Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio inform…