2 papers
cs.SD2025
On the de-duplication of the Lakh MIDI dataset
Eunjin Choi, Hyerin Kim, Jiwoo Ryu +2
A large-scale dataset is essential for training a well-generalized deep-learning model. Most such datasets are collected via scraping from various internet sources, inevitably intr…
cs.SD2024
Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
HaeJun Yoo, Hao-Wen Dong, Jongmin Jung +1
Representing symbolic music with compound tokens, where each token consists of several different sub-tokens representing a distinct musical feature or attribute, offers the advanta…