2 papers
eess.AS2026
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
Hokuto Munakata, Takehiro Imamura, Taichi Nishimura +1
We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no est…
cs.SD2025
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
Takehiro Imamura, Yuka Hashizume, Wen-Chin Huang +1
This paper proposes music similarity representation learning (MSRL) based on individual instrument sounds (InMSRL) utilizing music source separation (MSS) and human preference with…