From the 1 of 23 linked papers with an AI index.
23 papers
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Yi Luo, Rongzhi Gu, Jixun Yao
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with highe…
Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction
Renmingyue Du, Jixun Yao, Qiuqiang Kong +1
The paper proposes a reconstruction‑based method using autoencoders to detect out‑of‑distribution vocoder samples by reconstructing WavLM acoustic features, with contrastive learni…
SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
Chunbo Hao, Ruibin Yuan, Jixun Yao +5
Music structure analysis (MSA) underpins music understanding and controllable generation, yet progress has been limited by small, inconsistent corpora. We present SongFormer, a sca…
Aligning Generative Speech Enhancement with Perceptual Feedback
Haoyang Li, Nana Hou, Yuchen Hu +6
Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives tha…
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
Guobin Ma, Yuxuan Xia, Jixun Yao +5
This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The…
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
Yu Pan, Xiongfei Wu, Yuguang Yang +3
Despite recent advances in speech-to-speech translation (S2ST), it remains difficult to achieve both high translation accuracy and practical flexibility. In this paper, we present…