From the 1 of 5 linked papers with an AI index.
5 papers
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
The paper introduces VocalRender, a system that can directly synthesize singing voices from musical scores—including lyrics, pitches, note values, and tempo—without needing separat…
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealist…
Spiking Vocos: An Energy-Efficient Neural Vocoder
Yukun Chen, Zhaoxi Mu, Andong Li +2
Despite the remarkable progress in the synthesis speed and fidelity of neural vocoders, their high energy consumption remains a critical barrier to practical deployment on computat…
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
Zhaoxi Mu, Rilin Chen, Andong Li +3
This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios…
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
Zhaoxi Mu, Xinyu Yang, Gang Wang
While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, inclu…