10 papers
CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang +7
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoi…
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
Yizhong Geng, Wenxin Fu, Kecan Mao +7
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integratio…
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
Zhaoyang Meng, Zhengyao Ma, Kecan Mao +2
Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, whic…
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Yizhong Geng, Yanliang Li, Jinghan Yang +4
Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However, their effectiveness in low-res…
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
Cong Wang, Yizhong Geng, Yuhua Wen +7
Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce…
Hello-Chat: Towards Realistic Social Audio Interactions
Yueran Hou, Peilei Jia, Zihan Sun +5
Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer fr…