8 citations · 10 across the 13 of their papers we have counts for
11 papers · 1 filter
What Are You Listening to? Temporal Music Grounding for Audio-to-Text Large Language Models
Kun Fang, Ziyu Wang, Ichiro Fujinaga
Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce…
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
Jingwei Zhao, Gus Xia, Ziyu Wang +1
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. I…
Music-JEPA: Learning a World Model of Sound from Action
Ziyu Wang, Kun Fang, Yann LeCun
Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction f…
Real-Time Language Model Jamming: A Case Study for Live Music Accompaniment Generation
Bowen Zheng, Andrew H. Yang, Jiaqi Ruan +5
Language models (LMs) have become one of the most prominent paradigms in modern generative modeling. While making them faster has been the main focus of real-time deployment, speed…
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
Lekai Qian, Haoyu Gu, Jingwei Zhao +1
Tokenizing music to fit the general framework of language models is a compelling challenge, especially considering the diverse symbolic structures in which music can be represented…
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
Xiaoyu Yi, Qi He, Gus Xia +1
In automatic music generation, a central challenge is to design controls that enable meaningful human-machine interaction. Existing systems often rely on extrinsic inputs such as t…