4 papers
Latent Fourier Transform
Mason Wang, Cheng-Zhi Anna Huang
We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencod…
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ã-Tsang, Amdo and Kham Speech Dataset Generation
Yutong Liu, Ziyue Zhang, Ban Ma-bao +7
Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ã-Tsang, Amdo, and Kham), limiting progress in speech modeling. To addre…
Stemphonic: All-at-once Flexible Multi-stem Music Generation
Shih-Lun Wu, Ge Zhu, Juan-Pablo Caceres +2
Music stem generation, the task of producing musically-synchronized and isolated instrument audio clips, offers the potential of greater user control and better alignment with musi…
MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models
Shih-Lun Wu, Yoon Kim, Dave Carlton +3
We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an LLM's text vocabulary to include MIDI…