activity
20242026
collaborators

9 papers

cs.SD2026

Interpretable and Perceptually-Aligned Music Similarity with Pretrained Embeddings

Arhan Vohra, Taketo Akama

Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific tr…

cs.SD2026

PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation

Jaekwon Im, Natalia Polouliakh, Taketo Akama

Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human danc…

cs.SD2026

Self-supervised restoration of singing voice degraded by pitch shifting using shallow diffusion

Yunyi Liu, Taketo Akama

Pitch shifting has been an essential feature in singing voice production. However, conventional signal processing approaches exhibit well known trade offs such as formant shifts an…

cs.SD2026

Towards Realistic Synthetic Data for Automatic Drum Transcription

Pierfrancesco Melucci, Paolo Merialdo, Taketo Akama

Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are…

q-bio.NC2025

Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening

Taketo Akama, Zhuohao Zhang, Tsukasa Nagashima +3

Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, r…

q-bio.NC2025

SSDLabeler: Realistic semi-synthetic data generation for multi-label artifact classification in EEG

Taketo Akama, Akima Connelly, Shun Minamikawa +1

EEG recordings are inherently contaminated by artifacts such as ocular, muscular, and environmental noise, which obscure neural activity and complicate preprocessing. Artifact clas…