2 papers
cs.SD2025
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Jaejun Lee, Kyogu Lee
In this paper, we propose Vo-Ve, a novel voice-vector embedding that captures speaker identity. Unlike conventional speaker embeddings, Vo-Ve is explainable, as it contains the pro…
cs.SD2024
Do Captioning Metrics Reflect Music Semantic Alignment?
Jinwoo Lee, Kyogu Lee
Music captioning has emerged as a promising task, fueled by the advent of advanced language generation models. However, the evaluation of music captioning relies heavily on traditi…