activity
20242026
collaborators

34 papers

cs.CL2026

Audio-Based Understanding of Audiobook Narration Appeal

Shahar Elisha, Mariano Beguerisse-Díaz, Emmanouil Benetos

Narration is central to the audiobook listening experience, shaping how listeners engage with and understand the content. This work explores how narration qualities shape an audiob…

cs.SD2026

Velocity Prediction in Automatic Guitar Transcription

Jackson Loth, Xavier Riley, Simon Dixon +1

Automatic Music Transcription (AMT) models have achieved a high level of success in polyphonic transcription of various instruments. Velocity, typically a measure of note intensity…

cs.SD2026

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

Yinghao Ma, Haiwen Xia, Hewei Gao +9

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we…

cs.SD2026

Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound

Liumeng Xue, Ziya Zhou, Jiahao Pan +20

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and gene…

cs.SD2026

Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation

Nelly Garcia, Aditya Bhattacharjee, Gabryel Mason-Williams +3

Sound design workflows frequently oscillate between time-consuming library searches and the complexity of procedural synthesis, with practitioners typically relying on disconnected…

cs.SD2026

Text2Score: Generating Sheet Music From Textual Prompts

Keshav Bhandari, Sungkyun Chang, Abhinaba Roy +4

Developing text-driven symbolic music generation models remains challenging due to the scarcity of aligned text-music datasets and the unreliability of automated captioning pipelin…