3 papers
cs.SD2025
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Jaejun Lee, Kyogu Lee
In this paper, we propose Vo-Ve, a novel voice-vector embedding that captures speaker identity. Unlike conventional speaker embeddings, Vo-Ve is explainable, as it contains the pro…
cs.SD2024
Do Captioning Metrics Reflect Music Semantic Alignment?
Jinwoo Lee, Kyogu Lee
Music captioning has emerged as a promising task, fueled by the advent of advanced language generation models. However, the evaluation of music captioning relies heavily on traditi…
cs.SD2023
Combinatorial music generation model with song structure graph analysis
Seonghyeon Go, Kyogu Lee
In this work, we propose a symbolic music generation model with the song structure graph analysis network. We construct a graph that uses information such as note sequence and inst…