9 papers
Interpretable and Perceptually-Aligned Music Similarity with Pretrained Embeddings
Arhan Vohra, Taketo Akama
Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific tr…
PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation
Jaekwon Im, Natalia Polouliakh, Taketo Akama
Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human danc…
Self-supervised restoration of singing voice degraded by pitch shifting using shallow diffusion
Yunyi Liu, Taketo Akama
Pitch shifting has been an essential feature in singing voice production. However, conventional signal processing approaches exhibit well known trade offs such as formant shifts an…
Towards Realistic Synthetic Data for Automatic Drum Transcription
Pierfrancesco Melucci, Paolo Merialdo, Taketo Akama
Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are…
Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening
Taketo Akama, Zhuohao Zhang, Tsukasa Nagashima +3
Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, r…
SSDLabeler: Realistic semi-synthetic data generation for multi-label artifact classification in EEG
Taketo Akama, Akima Connelly, Shun Minamikawa +1
EEG recordings are inherently contaminated by artifacts such as ocular, muscular, and environmental noise, which obscure neural activity and complicate preprocessing. Artifact clas…