collaborators

9 papers

cs.SD2026

CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang +7

Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoi…

cs.SD2026

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

Yizhong Geng, Wenxin Fu, Kecan Mao +7

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integratio…

cs.SD2026

AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ

Zhaoyang Meng, Zhengyao Ma, Kecan Mao +2

Neural speech codecs have become the discrete interface between raw audio and speech language models, yet they remain optimized primarily for acoustic reconstruction fidelity, whic…

cs.CL2026

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Yizhong Geng, Yanliang Li, Jinghan Yang +4

Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However, their effectiveness in low-res…

cs.SD2026

Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention

Cong Wang, Yizhong Geng, Yuhua Wen +7

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce…

cs.SD2026

Hello-Chat: Towards Realistic Social Audio Interactions

Yueran Hou, Peilei Jia, Zihan Sun +5

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer fr…